Video compression system and method for reducing effects of packet loss over communication channel
Abstract
Problem to be solved.To provide a system and a method for reducing the influence of packet loss in a video communication system. A computer implementation method logically subdivides each of a series of images in a video stream into a plurality of tiles, each tile having a defined position within each of the series of images, and each piece of data. It involves packing tiles into multiple data packets and sending the data packets containing the tiles from the server to the client over the communication channel so as to maximize the number of tiles aligned at the packet boundaries. [Selection diagram] Fig. 10a

Term
10.1 yearsto projected expiry
Projected expiry 31 October 2036, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
2 claims: 2 independent, 0 dependent
- 1ビデオストリーム内のパケットロスの影響を減少するためのコンピュータ実施方法において、 ビデオストリームの一連の映像各々を複数のタイルへと論理的に細分化するステップであって、各タイルは、一連の各映像内に定義された位置を有するものであるステップと、 各データパケットの境界に整列されるタイルの数を最大にするように、タイルを複数のデータパケットへとパッキングするステップと、 タイルを含むデータパケットをサーバーから通信チャンネルを経てクライアントへ送信するステップと、を備えた方法。
- 2ビデオストリーム内のパケットロスの影響を減少するためのコンピュータ実施方法において、 ビデオストリームの一連の映像各々を複数のタイルへと論理的に細分化するステップであって、各タイルは、一連の各映像内に定義された位置を有するものであるステップと、 連続するデータパケットの境界間に延びるタイルの数を最小にするように、タイルを複数のデータパケットへとパッキングするステップと、 タイルを含むデータパケットをサーバーから通信チャンネルを経てクライアントへ送信するステップと、を備えた方法。
Independent claims2
305 paragraphs, as filed
0001The present invention generally relates to the field of data processing systems that improve the user's ability to operate and access audio and video media.
0002Related Application: This application is a partial continuation (CIP) application of Dec. 10, 2002, No. 10 / 315,460, entitled "APPARATUS AND METHOD FOR WIRELESS VIDEO GAMING", which was assigned to the assignee of this application. ..
0003Recorded audio and video media have represented the world since the days of Thomas Edison. To the beginning of the 20th century, recorded voice menu media (cylinders and records) and video media (Nickelodeon and movie), but has been distributed widely, both the technology was still young season. In the late 1920s, video was combined with audio on a mass market basis, and then became color video with audio. Radio broadcasting has gradually evolved into a large-scale advertising support form of broadcasting mass market audio media. With the establishment of television (TV) broadcasting standards in the mid-1940s, television joined radio as a form of broadcast mass-market media that delivers pre-recorded or live video to the home.
0004Until the mid-20th century, most U.S. households had gramophone record players to play recorded audio media, radios to receive live audio, and television receivers to show live audio / video (A / V) media. Was. Often, these three "media players" (record players, radios and TVs) were combined into a single cabinet that shared a common speaker, forming a home "media center." Media options were limited to consumers, but the media "ecosystem" was fairly stable. Most consumers knew how to use "media players" and were able to fully enjoy their capabilities. At the same time, media publishers (mainly video and television studios, as well as music companies) can sell their media both in theaters and at home without being bothered by popular pirated or "secondary sales" or resale of used media. I was able to distribute it to. Typically, publishers do not earn revenue from secondary sales, thus reducing the revenue that publishers may otherwise earn from second-hand media buyers for new sales. There are certainly used records sold in the middle of the 20th century, but such sales did not have a significant impact on record publishers. This is because music tracks are heard hundreds or thousands of times, unlike videos or video programs that adults typically watch only once or several times. Therefore, music media is far less "corrupted" than video / video media (ie, continues to be of value to adult consumers). When you buy a record, if the consumer likes the music, the consumer will probably hold it for a long time.
0005From the middle of the 20th century to the present day, the media ecosystem has undergone a series of fundamental changes in both the interests and losses of consumers and publishers. The widespread adoption of audio recorders, especially cassette tapes with high-quality stereo sound, certainly provided a high degree of consumer convenience. But it also marked the beginning of what is now widespread with the consumer media: pirated editions. Indeed, many consumers used cassette tapes to tape their own records purely for convenience, but increasingly more consumers (eg, dormitories with quick access to each other's record collections). Student) made a pirated copy. Consumers also taped music played over the radio, rather than buying records and tapes from publishers.
0006The advent of consumer VCRs can be set to record TV shows with new VCRs and can be viewed later, leading to further consumer convenience and "on-demand" based on movies and TV shows. It also led to the creation of a video rental business that can be accessed at. The rapid development of mass-market home media equipment since the mid-1980s has led to an unprecedented level of choice and consumer convenience, as well as a rapid expansion of the media publishing market.
0007Today, consumers are faced with excessive media choices and excessive media equipment, many of which are tied to a particular format of media or a particular publisher. Avid media consumers have stacked TVs and devices connected to computers in various rooms of the home, resulting in a "cable to one or more TV receivers and / or personal computers (PCs)". There is a "mouse burrow", as well as a group of remote controllers. (For this application, the term "personal computer" or "PC" refers to desktops, Macintosh® or other non-Windows® computers, Windows compatible devices, Unix® variations, laptops, etc. Refers to any type of computer suitable for home or office use, including video game consoles, VCRs, DVD players, audio surround-sound processors / amplifiers, satellite set top boxes, Includes cable TV set top box, etc. And for enthusiastic consumers, there are multiple similar functional devices due to compatibility issues. For example, consumers own both HD-DVD and Blu-ray DVD players, or Microsoft Xbox® and Sony. You may own both Playstation® video game systems. In fact, because some games are incompatible across game console versions, consumers may own both the XBox and subsequent versions, such as the Xbox 360®. Consumers are often confused about which video input and which remote controller to use. Even after the disc is placed in the correct player (eg DVD, HD-DVD, Blu-ray, Xbox or Playstation), video and audio inputs are selected for the device, and the correct remote controller is found, the consumer Still faces technical challenges. For example, in the case of widescreen DVDs, the user first determines the correct aspect ratio on his TV or monitor screen (eg 4: 3, full, zoom, wide zoom, cinema wide, etc.) and then sets. It is necessary to do. Similarly, the user will need to first determine the correct audio surround sound system format (eg AC-3, Dolby Digital, DTS, etc.) and then set it. Often, consumers are unaware that they are not enjoying the media content at the full capacity of their television or audio system (eg, watching a movie crushed with the wrong aspect ratio, or not surround sound. I listen to audio in stereo). 3. Full, zoom, wide zoom, cinema wide, etc.) must be determined and then set. Similarly, the user will need to first determine the correct audio surround sound system format (eg AC-3, Dolby Digital, DTS, etc.) and then set it. Often, consumers are unaware that they are not enjoying the media content at the full capacity of their television or audio system (eg, watching a movie crushed with the wrong aspect ratio, or not surround sound. I listen to audio in stereo). 3. Full, zoom, wide zoom, cinema wide, etc.) must be determined and then set. Similarly, the user will need to first determine the correct audio surround sound system format (eg AC-3, Dolby Digital, DTS, etc.) and then set it. Often, consumers are unaware that they are not enjoying the media content at the full capacity of their television or audio system (eg, watching a movie crushed with the wrong aspect ratio, or not surround sound. I listen to audio in stereo).
0008More and more Internet-based media devices are being added to the stack of devices. Audio devices such as the Sonos® Digital Music System stream audio directly from the Internet. Similarly, devices such as the Slingbox® entertainment player can record video and stream it over a home network or the Internet and view it on a PC at a remote location. Internet Protocol Television (IPTV) services provide cable TV-type services through Digital Subscriber Line (DSL) or other home Internet connections. Recently, Moxi® Media Center and Windows Efforts are also being made to integrate multiple media functions into a single device, such as a PC running XP Media Center Edition. Each of these devices provides an element of convenience for the function it performs, but lacks ubiquity and easy access to most media. In addition, such devices often require expensive processing and / or local storage, resulting in manufacturing costs often reaching hundreds of dollars. Moreover, these modern consumer electronics typically consume large amounts of power even while idle, which means that they are costly and waste energy resources over time. .. For example, a device may continue to operate if the consumer neglects to turn it off or switches to a different video input. And since none of the devices are a perfect solution, they have to be integrated with other stacks of devices in the home, which still leaves the user in the rat burrow of the wire and the sea of remote controls. It will be.
0009Moreover, many new Internet-based devices typically provide media in a more general form than it is available, when it works properly. For example, devices that stream video over the Internet often stream only video material, rather than the two-way "extra" that often accompanies a DVD, such as a video, game "production" or director's commentary. This is often due to the generation of bidirectional material in a particular format intended for a particular device that handles interactivity locally. For example, DVD, HD-DVD and Blu-ray discs have their own particular bidirectional format. Home media devices or local computers developed to support all popular formats probably require a level of sophistication and flexibility that makes them exorbitantly expensive and complex for consumers to operate. It will be.
0010In addition to this problem, if a new format is introduced in the future, the local device may not have the hardware capability to support the new format, which is consumer upgradeable. Means that you have to buy a local media device. For example, if high-definition video or stereo video (eg, one video stream for each eye) is introduced at a later date, the local device may not have the computing power to decode that video, or it may be new. You may not have the hardware to output the video in a different format (for example, if 60fps is given to each eye and the 120fps video is synced with the shuttered glasses to achieve a stereo feel. Assuming that if the consumer's video hardware can only support 60fps video, this option is not available without purchasing upgraded hardware).
0011The problem of aging and complexity of media equipment becomes a serious problem when it comes to sophisticated interactive media, especially video games.
0012Modern video game applications are primarily four major non-portable hardware platforms: Sony Playstation® 1, 2 and 3 (PS1, PS2 and PS3), Microsoft Xbox® and Xbox 360. (Registered Trademark), Nintendo Gamecube® (Registered Trademark) and Wii<sup>TM</sup>, As well as PC-based games. Each of these platforms, unlike the others, does not run games written to run on one platform on another. There is also the issue of compatibility from one generation of equipment to the next. Even if most software game developers create software games that are designed independently of a particular platform, in order to run a particular game on a particular platform, the game should be used on that particular platform. Requires a proprietary layer of software (often referred to as the "game development engine") to adapt. Each platform is sold to consumers as a "console" (ie, a stand-alone box attached to a TV or monitor / speaker), or the PC itself. Video games are typically Blu-ray containing video games that are implemented as sophisticated real-time software applications. Sold on optical media such as DVDs, DVD-ROMs, or CD-ROMs. As the speed of home broadband increases, video games are increasingly being used for download.
0013The specific requirements for platform compatibility with video game software are extremely stringent due to the real-time and computational requirements of advanced video games. For example, from one generation of video games to the next (eg, Microsoft Word), as well as general compatibility of manufacturing applications (eg, Microsoft Word) from one PC to another with a faster processing unit or core. , XBox to XBox360, or Playstation2 (PS2) to Playstation3 (PS3)) is expected to be fully game compatible. However, this is not the case with video games. Many written for previous generation systems, as video game manufacturers typically want the best possible performance for a given price point when a video game generation goes on sale. There are often abrupt architectural changes to the system that make the game not work on later generation systems. For example, the XBox is based on the x86 family of processors, while the XBox 360 is based on the PowerPC family.
0014Techniques for emulating traditional architectures are available, but given that video games are real-time applications, it is often impossible to achieve exactly the same behavior in emulation. This is a detriment to consumers, video game console manufacturers, and video game software publishers. For consumers, this means that old and new generations of video game consoles need to remain connected to the TV in order to be able to play all games. For console manufacturers, this means the costs associated with emulation and the hassle of adopting new consoles. Also, publishers need to launch multiple versions of new games to reach all potential consumers, i.e. launch versions for each video game brand (eg XBox, Playstation). In addition to doing so, it is often necessary to release a version for each given brand version (eg PS2 and PS3). For example, among other platforms, XBox, XBox Individual versions of Electronic Arts "Madden NFL 08" have been developed for 360, PS2, PS3, Gamecube, Wii and PC.
0015Portable devices such as cellular phones and portable media players also present a challenge to game developers. Gradually, such devices are connected to wireless data networks and are able to download video games. However, there are a wide variety of cell telephones and media devices on the market with a wide range of different display resolutions and computing powers. Also, such devices are typically constrained in power consumption, cost and weight, so advanced graphics such as graphics processing units (GPUs) such as devices manufactured by NVIDIA in Santa Clara, Calif. It lacks acceleration hardware. As a result, game software developers typically develop a given game title simultaneously for a number of different types of portable devices. The user finds that a given game title cannot be obtained for his or her particular cell phone or portable media player.
0016In the case of home video game consoles, hardware platform manufacturers typically impose royalties on software game developers for their ability to publish games on their platforms. Cell phone telegraph companies also typically impose royalties on game publishers to download games to cell phones. In the case of PC games, there is no royalties paid to publish the game, but game developers typically have a consumer service burden to support a wide range of possible PC configuration and installation issues. Due to its large size, it faces high costs. Also, PCs typically have few barriers to game software piracy. Because they are easy to reprogram by technically knowledgeable users, they can easily create pirated versions of the game and distribute them (eg, through the internet). Therefore, for software game developers, it is costly and disadvantageous to publish in game consoles, cell phones and PCs.
0017For console and PC software game publishers, the cost doesn't end there. In order to distribute the game through the retail channel, the issuer imposes a wholesale price lower than the retailer's selling price to earn a margin. Also, the publisher typically has to pay the manufacturing costs and distribute the physical media that holds the game. The issuer may also expect, for example, if the game does not sell, if the price of the game drops, or if the retailer has to refund some or all of the wholesale price and / or take the game from the buyer. "Price protection costs" are often imposed by retailers to cover this situation. In addition, retailers typically charge publishers to help bring games to market with advertising leaflets. In addition, retailers are increasingly buying back games from users who have finished playing them, selling them as used games, and typically not sharing the revenue of used games with game publishers. In addition to the cost burden imposed on game publishers, pirated versions of games are often created and distributed over the Internet for users to download and copy for free.
0018Games are becoming more common as internet broadband speeds are increasing, and broadband connections to homes and internet "cafes" where internet-connected PCs are rented are becoming more common in the United States and around the world. , Is being distributed more and more after being downloaded to a PC or console. Also, broadband connections are increasingly being used to play multiplayer and large-scale multiplayer online games (both of which are referred to by the acronym "MMOG" in this disclosure). These changes reduce some of the issues and costs associated with retail distribution. Downloading online games presents some drawbacks to game publishers in that distribution costs are typically low and there is little or no cost from unsold media. However, downloaded games are still pirated, and due to their size (often a few gigabytes in size), it takes a very long time to download. In addition, small disk devices such as those sold with portable computers or video game consoles are filled with multiple games. However, the problem of pirated editions has been mitigated to the extent that games or MMOGs require an online connection for playable games. This is because the user is usually required to have a valid user account. Unlike linear media (eg video and music) that can be copied by a microphone that records audio from a camera or speaker that captures video on the display screen, each video game experience is unique and simple video / audio. Cannot be copied using recording. Therefore, even in areas where copyright law is not strongly enforced and piracy is rampant, MMOGs can be shielded from piracy and therefore support their business. For example, Vivendi SA's World of
0019While online nature or MMOGs can often mitigate piracy, online game operators still face the remaining challenges. Many games require substantial local (ie, home) processing resources for online or MMOG to function properly. Users may not be able to play games if they have a poorly performing local computer (eg, a low-end laptop that does not have a GPU). Moreover, as game consoles get older, they gradually retreat from the latest and become unable to handle more advanced games. There is often installation complexity, even assuming that the user's local PC can handle the game's compute requests. The driver may be incompatible (for example, when a new game is downloaded, a new version of the graphics driver will be installed, which will not match previously installed games based on the older version of the graphics driver. May be activated). The console runs out of local disk space as more games are downloaded. Complex games are typically found by game developers when bugs are found and repaired, or when changes are made to the game (for example, the level of the game is too difficult or too easy to play). If you do), you will receive a patch downloaded from the game developer over time. These patches require new downloads. However, sometimes not all users complete the download of all patches. At other times, downloaded patches introduce other compatibility or disk space consumption issues.
0020Also, during game play, large data downloads may be required to provide graphics or behavior information to the local PC or console. For example, if a user enters an MMOG room and encounters a scene or character that is created with graphic data or has behavior that is not available on the user's local machine, the scene or character data must be downloaded. It doesn't become. As a result, if the internet connection is not fast enough, there will be a substantial delay during game play. And if the scene or character encountered demands storage space or computing power that goes beyond the local PC or console, the user may not be able to proceed with the game or must continue with poor quality graphics. Therefore, online or MMOG games often limit their memory and / or computational complexity requirements. Moreover, they often limit the amount of data transfer during the game. Online or MMOG may also narrow the market for users who can play games.
0021In addition, more and more technically knowledgeable users are increasingly modifying their games to reverse engineer and cheat local copies of the game. Cheating is as easy as repeating a button press faster than humans can (for example, shooting a gun very fast). In games that support in-game asset transactions, cheats reach a level of sophistication that results in fraudulent transactions involving assets of real economic value. When the online or MMOG economic model is based on such asset transactions, this has substantially detrimental consequences for game operators.
0022The cost of developing new games is that PCs and consoles can create increasingly sophisticated games (with more realistic graphics such as real-time ray tracing and more realistic behavior such as real-time physics simulation). It is increasing with. In the early days of the video game industry, video game development was a process very similar to application software development, i.e. most of the development cost was software development, graphics, audio, and behavioral elements or "assets". It is not developed for videos with a wide range of special effects, for example. Today, many elaborate video game development efforts are more closely like special effects video development than software development. For example, many video games give simulations of the 3D world and generate more and more hospitalistic (ie, computer graphics that look as realistic as live motion video photography) characters, props and environments. To do. One of the most challenging aspects of thermalistic game development is to create a computer-generated human face that is indistinguishable from a living human face. Contour developed by Mova in San Francisco, California<sup>TM</sup>Face capture techniques such as reality capture capture and track the exact shape of the performer's face with high resolution during exercise. This technology allows 3D faces that are virtually indistinguishable from captured live-moving faces to be rendered on a PC or game console. Capturing and rendering "photoreal" human faces is useful in many ways. First, very recognizable celebrities or sports players (often hired at high cost) are often used in video games, where imperfections are obvious to the user and confuse the viewing experience. May cause discomfort or discomfort. Also, in many cases, achieving a high degree of photorealism requires a high degree of detail, i.e., potentially allowing the polygons and / or textures to change from frame to frame as the face moves. It is required to render a large number of polygons and high resolution textures.
0023When a large number of polygonal scenes with detailed textures change rapidly, the PC or game console that supports the game will provide sufficient polygon and texture data for the required number of animation frames generated in the game segment. May not have enough RAM to store. In addition, a single optical drive or single disk drive typically available for PCs or consoles is typically much slower than RAM, and typically GPUs when rendering polygons and textures. Cannot maintain the maximum acceptable data rate. Current games typically load most of the polygons and textures into RAM, which means that the complexity and time width of a given scene is largely limited by the amount of RAM. For example, in the case of face animation, this would bring the PC or game console to a non-photoreal low resolution before the game pauses and loads polygons and textures (and other data) for more frames. Limited to either a face or a polygonal face that can only be animated for a limited number of frames.
0024Watching the progress bar slowly move across the screen as the PC or console displays a message similar to Loading ... is a drawback inherent to today's users of complex video games. It is accepted. Delay between loading the next scene from a disk (where "disk" refers to non-volatile optical or magnetic media, as well as non-disk media, such as semiconductor "flash" memory, unless otherwise indicated). Takes seconds or minutes. This is a waste of time and can be quite frustrating for game players. As mentioned above, many or all delays are due to the time it takes to load polygons, textures or other data from the disk, but the processor and / or GPU in the PC or console takes the data for the scene. Part of the loading time may be spent during preparation. For example, a soccer video game allows a player to choose from a large number of players, teams, stadiums and weather conditions. Therefore, different polygons, textures and other data (collectively "objects") are required for the scene based on what particular combination is selected (eg, different teams are in their uniform). Has different colors and patterns). Many or all various permutations can be enumerated, many or all objects can be pre-computed, and those objects can be stored on the disk used to store the game. However, if the number of permutations is large, the amount of storage required for all objects may be too large to fit on an disk (impossible to download). Therefore, existing PC and console systems are typically constrained by both the complexity of a given scene and the play time width, and suffer from long load times for complex scenes.
0025Another notable limitation with traditional video game systems and application software systems is the large database of 3D objects such as polygons and textures that need to be loaded for processing into a PC or game console, for example. Is gradually being used. As mentioned above, such databases require long load times when stored locally on disk. However, load times are usually even more stringent when the database is stored in a remote location and accessed over the Internet. In such a situation, it may take minutes, hours, or even days to download a large database. In addition, such databases are often generated at great cost (eg, 3D models of sailboats with detailed high masts for use in games, movies or historical documentaries) and sold to local end users. Intended to. However, databases can be pirated when downloaded to local users. In many cases, the user simply evaluates the database to see if it suits the user's needs (eg, if the 3D costume of the game character has a satisfactory look or appearance when the user makes certain movements). I would like to download the database for. Long load times hinder users who evaluate 3D databases before deciding to buy.
0026MMOGs have similar problems, especially as games that allow users to gradually use customized characters. In the case of a PC or game console for displaying a character, access to a database of 3D geometry (polygons, textures, etc.) and behavior towards that character (eg, if the character has a shield, the shield is a spear. It is necessary to obtain (whether it is strong enough to distract). Typically, when the MMOG is first played by the user, a large database of characters is pre-available with an initial copy of the game, which is either locally available on the game's optical disc or downloaded to the disc. To. However, as the game progresses, if a user encounters a character or object whose database is not locally available (for example, if another user creates a customized character), that character or object will be displayed. By the time you can, the database must be downloaded. This causes a substantial delay in the game.
0027Given the sophistication and complexity of video games, another challenge for video game developers and publishers on traditional video game consoles is often a couple of years at the cost of tens of millions of dollars. It is to develop a video game by multiplying. Assuming that a new video game console platform is introduced approximately once every five years, game developers will be able to get video games as soon as the new platform is launched. Development work needs to start just these game years before the launch of the new game console. A large number of consoles are sometimes released by competing manufacturers at about the same time (eg, within a year or two of each other), but what you have to look at is the popularity of each console, for example, which console Will it generate the maximum sales of video game software? For example, in recent console cycles, Microsoft XBox360, Sony Playstation 3 and Nintendo The Wii is scheduled to be deployed in about the same general time frame. But years before its introduction, game developers must essentially "bet" on which console platform will be more successful than others, and dedicate development resources accordingly. Video production companies must also allocate limited production resources based on estimates of the likelihood of a movie's success well before the movie's release. Given the increasing investment levels required for video games, game production is becoming more and more similar to video production, and game production companies routinely devote production resources to estimates of the future success of a particular video game. Dedicated to the target. However, unlike video companies, this bet is not simply based on the success of the production itself, but rather on the success of the game console intended to run the game. You can mitigate the risk by launching the game for multiple consoles at once, but this additional effort increases costs and often delays the actual launch of the game.
0028Application software and user environments on PCs are computationally more central, dynamic and interactive in order to be more effective and intuitive as well as more visually appealing to the user. It becomes. For example, the new Windows Vista<sup>TM</sup>The operating system, and successive versions of the Macintosh® operating system, combine visual animation effects. Maya from Autodesk<sup>TM</sup>Advanced graphics tools such as provide highly sophisticated 3D rendering and animation capabilities that push the boundaries of modern CPUs and GPUs. However, the computational requirements of these new tools pose a number of practical problems for users and software developers of such products.
0029The OS graphics requirement is that the visual display of the operating system (OS) must work on a wide variety of computers, including previous generation computers that are no longer sold but can still be upgraded with newer OS. , Typically largely limited by the minimum denominator of the computer that is the target of the OS, including non-GPU computers. This severely limits the graphics capabilities of the OS. In addition, battery-powered portable computers (eg, laptops) limit their visual display capabilities. This is because high computational activity on the CPU or GPU typically increases power consumption and shortens battery life. Portable computers typically include software that automatically reduces processor activity to reduce power consumption when the processor is not in use. In some computer models, users can manually reduce processor activity. For example, Sony's VGN-SZ280P laptop has "Stamina" on one side (for low performance and long battery life) and "Speed" on the other (for high performance and short battery life). Includes switches that have been turned on. An operating system running on a portable computer must be able to function usefully, even if the computer runs at part of its peak performance capabilities. Therefore, OS graphics performance often remains well below the latest computing power available.
0030High-end compute-intensive applications like Maya are often sold with the expectation that they will be used on high-performance PCs. This typically establishes the lowest common denominator requirement of very high performance and more expensive and less portable. As a result, such applications are a generic OS (or Microsoft). It has a much more limited target audience than general purpose production applications such as Office) and is typically sold in significantly smaller quantities than general purpose OS software or general purpose application software. Also, potential viewers are further limited, as it is often difficult for prospective users to try out such computationally-friendly applications in advance. For example, Maya wants students to learn how to use Maya, or before potential buyers who are already knowledgeable about such applications invest in purchases (including buying high-end computers that can run Maya). Suppose you want to try. Students or potential buyers can download a demo version of Maya or get a physical media copy of it, but they don't have a computer that can run Maya at its full potential (eg, dealing with complex 3D scenes). In some cases, a complete informal assessment of the product cannot be performed. This limits the viewership of such high-end applications. It also contributes to high selling prices. This is because development costs are usually allocated over a much smaller number of purchases than general purpose applications.
0031High-priced applications also create a great incentive for individuals and businesses to use pirated copies of application software. As a result, high-end application software is plagued by pirated versions, despite significant efforts by publishers of such software to mitigate piracy through a variety of technologies. Even when using pirated high-end applications, users cannot rule out the need to invest in expensive, modern PCs to perform pirated copies. Thus, users of pirated software can use the software application for a portion of its actual retail price, but are still required to purchase or obtain an expensive PC in order to take full advantage of the application.
0032The same is true for users of high-performance pirated video games. Pirates can get the game for a portion of their actual price, but the expensive computing hardware needed to play the game properly (eg, GPU-enhanced PC, or XBox 360, etc.) High-end video game consoles) are still required to be purchased. Assuming that video games are typically a pastime for consumers, the additional costs for high-end video game systems are exorbitant. This is even worse in countries where the average annual salary of workers today is significantly lower than in the United States (eg China). As a result, only a very small proportion of the population owns high-end video game systems or high-end PCs. In these countries, "Internet cafes" where users pay for the use of computers connected to the Internet have become quite common. In many cases, such Internet cafes only have older models or low-end PCs that do not have high-performance features such as GPUs that allow players to play computationally-friendly video games. This is a key factor in the success of games that run on low-end PCs, such as Vivendi's "World of Warcraft," which is very successful in China and is commonly played in Internet cafes there. In contrast, computationally-friendly games like "Second Life" are extremely unlikely to be playable on PCs installed in Chinese internet cafes. Such games are virtually inaccessible to users who can only access low-performance PCs in Internet cafes.
0033There is also a barrier for users who are thinking about buying a video game and want to try it by first downloading a demo version of the game to their home via the internet. Video game demos are often full-fledged versions of a game in which some features are disabled or the amount of play in the game is limited. This can involve a long process (perhaps hours) of downloading gigabytes of games before they can be installed and run on either a PC or console. In the case of a PC, can the special drivers required for the game (eg DirectX or OpenGL drivers) be calculated, the correct version downloaded, installed, and then the PC can play the game? Including deciding whether. This latter step is for the PC to have enough processing power (CPU and GPU), enough RAM, and a compatible OS (eg, some games are Windows). Includes determining if you have (runs on XP but not on Vista). Therefore, after a long process of trying to run a video game demo, the user can fully find that the video game demo is probably unplayable, given the user's PC configuration. Worse, when users download new drivers to try out the demo, those driver versions may not be compatible with other games or applications that users normally use on their PCs, and therefore , Installing the demo may cause previously working games or applications to fail. These barriers not only frustrate users, but also barrier video game software publishers and video game developers who bring games to market.
0034Another issue that leads to financial inefficiencies is that a given PC or game console is typically designed to accept a certain level of performance requirements for applications and / or games. For example, some PCs have some RAM, slow or fast CPUs, and if they have GPUs, slow or fast GPUs. Some games or applications take advantage of the total computing power of a given PC or console, but many games or applications do not. If the user's choice of game or application does not reach the peak performance capabilities of the local PC or console, the user will waste money on the PC or console with respect to unused features. In the case of consoles, the console manufacturer has paid more than was needed to subsidize the console costs.
0035Another problem that exists when buying and enjoying a video game is to allow the user to see someone else playing the game before the user decides to buy the game. There are numerous traditional solutions for recording video games for later playback. For example, US Pat. No. 5,558,339 teaches that game information, including game controller actions, is recorded during "gameplay" on a video game client computer (owned by the same or different users). There is. This state information can later be used to play some or all game actions on a video game client computer (eg, a PC or console). A notable drawback of this solution is that in order for the user to see the recorded game, the user owns a video game client computer that can play the game and the gameplay is played when the recorded game state is played. Must have a video game application running on that computer to be identical. Besides this, the video game application must be written so that there is no difference in execution between the recorded game and the played game.
0036For example, game graphics are generally calculated frame by frame. For many games, the game logic is sometimes another process of removing the CPU cycle from the game application on a PC, whether the scene is particularly complex or has other delays that slow it down (for example, on a PC). It takes less than one frame or longer to calculate the graphic displayed for the next frame, based on whether it is executed). In such games, a calculated "threshold" frame will eventually occur that is slightly less than one frame time (eg, a few CPU clock cycles shorter). When that same scene is recalculated using the exact same game state information, it can easily take several CPU clock cycles longer than one frame time (for example, the internal CPU bus is slightly out of phase with the external DRAM bus). And if you want to introduce a time delay of a few CPU cycles, even without the big delay of another process that takes a few milliseconds of CPU time from the game processing). Therefore, when the game is played, the frames are calculated in two frame times instead of one frame time. Some behaviors are based on how often the game calculates new frames (eg, when the game samples input from the game controller). While the game is displayed, this difference in time criteria for different behavior does not affect game play, but results in different results for the played game. For example, the basketball trajectory is calculated at a constant 60fps velocity, but if the game controller input is sampled based on the calculated frame velocity, the calculated frame velocity will be recorded by the game. When played, it's 53fps, but when the game is played, it's 52fps, which makes a difference whether basketball is blocked from entering the basket and has different consequences. Therefore, use the game state to video game
0037Another conventional solution for recording video games is simply recording the video output of a PC or video game system (eg, to a VCR, DVD recorder, or to a video capture board on a PC). The video can then be rewound and played, or the recorded video can typically be compressed and then uploaded to the Internet. The drawback of this solution is that when the 3D game sequence is played, the user is limited to viewing the sequence only from the point of view in which the sequence was recorded. In other words, the user cannot change the viewpoint of the scene.
0038In addition, when a compressed video of a recorded game sequence played on a home PC or game console is made available to other users over the Internet, the compressed video will be compressed in real time, even if the video is compressed in real time. It is not possible to upload to the internet. The reason is that many homes around the world connected to the Internet have extremely asymmetric broadband connections (for example, DSL and cable modems typically have a much wider downstream bandwidth than the upstream bandwidth. It has a width). Compressed high-resolution video sequences often have more bandwidth than the network's upstream bandwidth capacity, making it impossible to upload them in real time. Therefore, after the game sequence is played, there is a significant delay (perhaps minutes or hours) before another user on the Internet can see the game. This delay is acceptable in some situations (eg, looking at previously occurring game player performance), but the ability to see the game live (eg, a basketball tournament played by a champion player), or the game playing live. Eliminate the "instantaneous playback" ability when done.
0039Another traditional solution is to allow viewers to watch the video game live on the television receiver only under the control of the television producer. Several channels in the United States and other countries offer video game viewing channels, and TV viewers watch some video game users (eg, top rated players playing in tournaments) on the video game channel. Can be done. This is achieved by feeding the video output of a video game system (PC and / or console) to a video distribution and processing device for television channels. This is different from when a television channel broadcasts a live basketball game in which many cameras send live footage around the basketball court from different angles. Thus, television channels can utilize their video / audio processing to act on the device to manipulate the output from various video game systems. For example, a TV channel can superimpose text indicating the state of different players on top of a video from a video game (as if it were overlay text in a live basketball game), and a TV channel can be a game. It can be recorded overlaid with audio from a commentator who can discuss the actions that occur inside. In addition, the video game output can be combined with a camera that records the video of the actual player in the game (eg, showing an emotional response to the game).
0040One problem with this solution is that such live video feeds must be available in real time to the video distribution and processing equipment of television channels in order to obtain the excitement of live broadcasting. However, as mentioned above, this is often not the case when the video game system is run from home, especially if part of the broadcast contains live video from a camera that captures the game player's real-world video. It is possible. Further, in the tournament state, as described above, there is also a problem that the game player in the home changes and cheats the game. For these reasons, such video game broadcasts over television channels have players and video game systems gathered in a common location (eg, a television studio or stadium), where television production equipment has a large number of video games. Often configured to accept video feeds from the system and potential live cameras.
0041Such traditional video game television channels have an experience similar to a live sporting event, with respect to both action in the video game world and action in the real world, for example, allowing the video game player to be portrayed as an "athlete". Although very exciting screenings can be given to television viewers, these video game systems are often limited to a state in which the players are physically very close to each other. And since the television channels were broadcast, each broadcast channel can only show one video stream selected by the producer of the television channel. Due to these limitations and the high cost of broadcast time, production equipment and producers, television channels typically only show the highest rated players to play in top tournaments.
0042Moreover, a given television channel that broadcasts a full-screen video of a video game to all television viewers will only show one video game at a time. This severely limits the choices of television viewers. For example, a television viewer may not be interested in a game that is shown at a given time. Another viewer is only interested in watching the gameplay of a particular player at a given time, which is not featured by the television channel. In other cases, the viewer is only interested in seeing how a professional player handles a particular level in the game. Yet other viewers want to control the perspective of watching the video game, which is different from the one selected by the production team or the like. In short, television viewers have a myriad of preferences for watching video games that are unacceptable by a particular broadcast on a television network, even when many different television channels are viewed. For all the reasons mentioned above, traditional video game TV channels have significant limitations in showing video games to TV viewers.
0043Another drawback of traditional video game systems and application software systems is that they are complex and usually plagued by errors, crashes and / or unintended and undesired behavior (collectively "bugs"). That is. Games and applications typically go through a debugging and tuning process (often referred to as "Software Quality Assurance" or SQA) before launch, but almost always the game or application is a widespread audience in the field. When it is released in, a bug suddenly appears. Unfortunately, it is difficult for software developers to identify and track many bugs after launch. It is difficult for software developers to notice bugs. Even when you learn about bugs, there is only a limited amount of information available to identify what caused the bug. For example, when a user calls the game developer's customer service line and plays a game, the screen starts flashing and then turns dark blue, leaving a message indicating that the PC freezes. This gives the SQA team very little information to help track down bugs. A game or application connected online can sometimes provide a lot of information in some cases. For example, a "watchdog" process can sometimes be used to monitor a game or application for a "crash." The watchdog process collects statistics about the state of a game or application process when it crashes (eg, stack state, memory usage, how far the game or application has progressed, etc.), and then collects that information. , Can be uploaded to the SQA team via the internet. However, in complex games or applications, such information can take a very long time to decipher in order to determine exactly what the user did in the event of a crash. Therefore, what kind of event sequence is
0044Yet another problem with PCs and game consoles is the problem of services that cause considerable inconvenience to consumers. Also, these service issues affect PC or game console manufacturers. This is because the manufacturer must send a special box to safely transport the broken PC or console and bear the cost of repair if the PC or console is under warranty. .. Publishers of game or application software are also affected by lost sales (or use of online services) due to the PC and / or console being in a repaired state.
0045Figure 1 shows Sony Playstation® 3, Microsoft XBox 360®, Nintendo Wii<sup>TM</sup>, Window-based personal computer, or Apple Indicates a traditional video game system such as the Macintosh. Each of these systems is a central processing unit (CPU) for executing program code, typically a graphics processing unit (GPU) for performing advanced graphics operations, and for communicating with external devices and users. It has multiple forms of input / output (I / O). For simplicity, these components are shown combined together as a single unit 100. Further, the conventional video game system of FIG. 1 plays an optical media drive 104 (for example, a DVD-ROM drive), a hard drive 103 for storing a video game program code and data, and a multiplayer game. , Network connection 105 for downloading patches, demos or other media, Random access memory (RAM) 101 for storing program code currently being executed by CPU / GPU 100, commands entered by the user during gameplay It is shown to include a game controller 106 for receiving, and a display device 102 (eg, SDTV / HDTV or computer monitor).
0046The traditional system shown in Figure 1 suffers from a number of limitations. First, the optical drive 104 and the hard drive 103 tend to have very slow access speeds compared to the RAM 101. When functioning directly through RAM 101, the CPU / GPU 100 can actually process far more polygons per second than is possible when program code and data are read directly from hard drive 103 or optical drive 104. .. This is because RAM101 generally has a very high bandwidth and is not plagued by the relatively long seek delay of the disk mechanism. However, these traditional systems provide only a limited amount of RAM (eg 256-512 Mbytes). Therefore, a Loading ... sequence is often required in which the RAM 101 is periodically filled with data for the next sequence of video games.
0047Some systems attempt to superimpose program code loading at the same time as gameplay, but this can only be done when there is a known sequence of events (for example, when driving a car down the road). You can load the geometry of an approaching building by the road while driving a car). For complex and / or rapid scene changes, this form of superposition usually does not work. For example, if the user is in a battle and the RAM101 is completely filled with data representing objects in the field of view at that moment, the user quickly left the field of view to see the objects that are not currently loaded in the field of view. When moved to, an action discontinuity occurs. This is because there is not enough time to load new objects from hard drive 103 or optical media 104 into RAM 101.
0048Another problem arises in the system of FIG. 1 due to the limited storage capacity of the hard drive 103 and the optical media 104. Although it is possible to manufacture a disk storage device with a relatively large storage capacity (for example, 50 GB or more), it is still not possible to provide sufficient storage capacity for some scenes currently encountered in video games. For example, as mentioned above, soccer video games allow users to choose from a large number of teams, players and stadiums around the world. For each team, each player and each stadium, a large number of texture and environment maps are required to characterize 3D curves around the world (eg, each team has its own jersey, each of which has its own unique jersey. Requires a unique texture map).
0049One technique used to address this latter issue is for the game to pre-calculate textures and environment maps when they are selected by the user. This may involve a number of computationally intensive processes, including video decompression, 3D mapping, shading, data structure organization, and so on. As a result, there may be a delay for the user while the video game performs these calculations. One way to reduce this delay is, in principle, to perform all of these calculations when the game is first developed, including the team, player roster and stadium permutations. Therefore, the released version of the game downloads all this preprocessed data stored on the optical media 104, or one or more servers on the internet, to the hard drive 103 over the internet when the user makes a selection. Includes with a given team, player roster, and selected preprocessed data for stadium selection. However, as a practical matter, such preloaded data in each permutation considered in gameplay will probably be terabytes of data that far exceeds the capacity of today's optical media devices. Moreover, the data for a given team, player roster, and stadium selection will probably be hundreds of megabytes or more. For example, with a 10 Mbps home network connection, downloading this data through the network connection 105 would take more time than calculating the data locally.
0050Therefore, the traditional game architecture shown in FIG. 1 causes the user to be significantly delayed during the transition to the main scene of a complex game.
0051Another problem with the traditional solution shown in Figure 1 is that video games tend to be more advanced year by year, requiring higher CPU / GPU power. Therefore, even assuming a limited amount of RAM, the requirements of video game hardware exceed the peak level of processing power available for these systems. As a result, users will need to upgrade their game hardware every few years (or play new games at a lower level) to keep pace. As a result of one of the constant advances in video games, video gameplay machines for home use are typically economical because their cost is usually determined by the requirements of the highest performance games that can be supported. Inefficient. For example, the XBox 360 is used to play games like the "Gears of War" that require a high-performance CPU, GPU, and hundreds of megabytes of RAM, or the XBox 360 is a few kilobytes of RAM. And used to play Pac Man, a 1970s game that requires only a very low performance CPU. In fact, the XBox 360 is a Pac It has enough computing power to host many Man games at once.
0052Video game machines are typically turned off most of the week. On average, active games spend 14 hours each week playing console video games, or one week, according to the July 2006 issue of the Nielsen Entertainment Study on active games over 13 years old. Only 12% of the total time. This means that on average, video game consoles are idle for 88% of the time, making expensive resources inefficient. Assuming that video game consoles are often subsidized by manufacturers to lower purchase prices (in the hope that the subsidies will be returned by royalties from future video game software purchases), this is This is especially significant.
0053Video game consoles also bear the costs associated with most consumer electronics. For example, the electronics and mechanisms of the system need to be housed in an enclosure. The manufacturer must provide a repair warranty. The retailer that sells the system needs to collect a margin on the sale of the system and / or the sale of the video game software. All of these factors add to the cost of the video game console, which must be subsidized by the manufacturer, passed on to the consumer, or both.
0054Moreover, pirated editions are a major issue for the video game industry. The security mechanisms used in virtually all major video game systems are "cracked" year after year, with unauthorized copying of video games. For example, the XBox 360 security system was cracked in July 2006, and users can now download unauthorized copies online. Downloadable games (eg games for PC or Mac) are particularly vulnerable to piracy. In some parts of the world where piracy is weakly cracked down, there are essentially no successful markets for standalone video gaming software. That's because users can buy a pirated copy as easily as a legitimate copy for a fraction of its cost. Also, in many parts of the world, the cost of game consoles is a high percentage of revenue, and cracking down on pirated games can only afford the latest gaming systems.
0055In addition, the used game market will reduce the revenue of the video game industry. When a user gets tired of a game, he or she can sell the game to a store, which resells the game to another user. This unauthorized but common practice significantly reduces the income of game publishers. Similarly, when there is a platform shift every few years, there is usually a 50% drop in sales. This means that when the user learns that a new version of the platform is about to go on sale, he or she will stop buying games for the older platform (for example, when Playstation 3 is about to go on sale, the user will buy a game on the Playstation 2. (Stop). The combined loss of sales associated with new platforms and increased development costs will have a tremendous impact on the interests of game developers.
0056Also, new game consoles are very expensive. The XBox 360, Nintendo Wii, and Sony Playstation 3 will all be retailed for hundreds of dollars. High-power personal computer gaming systems cost up to $ 8000. This represents a significant investment for users, especially given that the hardware will be obsolete after a few years and many systems will be purchased for children.
0057One solution to these problems is an online game in which the game program code and data are hosted on a server and delivered on demand as compressed video and audio streamed over a digital broadband network to client machines. Some companies, such as Finland's G-Cluster (now a subsidiary of Softbank Broadmedia in Japan), offer these services online. Similar gaming services are available on local networks, such as in hotels, and are provided by DSL and cable television providers. The main drawback of these systems is the latency, that is, the time it takes for the signal to travel to and from a game server typically located at the operator's "headend." High-speed action video games (also known as "twitch" video games) are very much between the time the user takes action on the game controller and the time the display screen is updated to show the result of the user action. Request a short waiting time. A short wait time is required to give the user the feeling that the game responds "immediately". The user can be satisfied with different waiting time intervals based on the format of the game and the skill level of the user. For example, for slow casual games (such as backgammon) or games that play the role of slow action, a wait time of 100 ms is acceptable, but for high speed action games, when the wait time exceeds 70 or 80 ms, the user , The performance is insufficient in the game and is unacceptable. For example, in a game that requires fast reaction time, the accuracy drops sharply as the wait time increases from 50 to 100 ms.
0058When the game or application server is installed in a nearby controlled network environment, or in an environment where the network path to the user is predictable and / or the bandwidth peak is acceptable, the maximum latency and It is much easier to control latency for both latency consistency (for example, as a user observes constant movement from a digital video streaming over a network). These levels of control are between the headend of the cable TV network and the home of the cable TV subscriber, or between the DSL central office and the home of the DSL subscriber, or local to the commercial office from the server or user. It can be achieved in an area network (LAN) environment. It is also possible to obtain special grade point-to-point private connections between companies with guaranteed bandwidth and latency. However, in a game or application system that hosts a game in a server center connected to the general Internet and then streams compressed video to the user via a broadband connection, many factors cause waiting time, and the conventional system is deployed. Causes tremendous restrictions.
0059In a typical broadband home, the user can have a DSL or cable modem for broadband services. Such broadband services usually incur a round-trip latency of about 25 ms (sometimes longer) between the user's home and the general Internet. In addition, there is a round-trip delay due to the routing of data to the server center via the Internet. The latency through the Internet varies based on the route on which the data is given and the delay incurred when it is routed. In addition to routing delays, there is also a round-trip delay due to the speed of light traveling through the fiber optics that interconnect most of the Internet. For example, every 1000 miles, due to the speed of light through the fiber optics and other overhead, it suffers a round trip latency of about 22 ms.
0060The data rate of the data streamed over the Internet creates additional latency. For example, the user says "6Mbps In fact, when having a DSL service sold as a "DSL service", the user probably gets less than 5 Mbps downstream throughput, and probably during peak load times in a digital subscriber line access multiplexer (DSLAM). Due to various factors such as congestion, connection quality degradation will be seen periodically. Also, if the locally shared coaxial cable that is looped through an adjacency or somewhere else in the cable modem system network becomes congested, the cable modem data used for the connection sold as the "6Mbps Cable Modem Service" A similar problem occurs that reduces the rate much lower. If a data packet at a constant rate of 4 Mbps is streamed from the server center via such a connection in one direction in the User Datagram Protocol (UDP) format, then if everything works, the data packet will be additional. Passing through without any latency, but in the presence of congestion (or other obstacles) and only 3.5 Mbps available to stream data to the user, packets are typically dropped. Either data loss occurs or the packet is queued to the congestion point until the packet can be sent, resulting in additional latency. Different congestion points have different queuing capacities to hold delayed packets, so in some cases packets that cannot pass through the congestion are dropped immediately. In other cases, a few megabits of data are queued and finally sent out. However, in almost all cases, the queue at the congestion point has a capacity limit, and if these limits are exceeded, the queue will overflow and packets will be dropped. Therefore, it is necessary not to exceed the data rate capacity from the game or application server to the user in order to avoid incurring additional latency (or, worse, packet loss).
0061It also suffers from latency due to the time required to compress the video on the server and decompress the video on the client device. In addition, the video game running on the server suffers a wait while calculating the next frame to be displayed. Currently available video compression algorithms suffer from either high data rates or long latency. For example, Motion JPEG is an intra-frame-only Rossy (intra frame-only) that features low latency. lossy) A compression algorithm. Each video frame is compressed independently of each other's video frames. When the client device receives a frame of compressed motion JPEG video, it immediately decompresses and displays the frame, resulting in very short latency. However, because each frame is compressed separately, the algorithm cannot take advantage of the similarities between successive frames, and as a result, the intraframe-only video compression algorithm suffers from very high data rates. ing. For example, 60fps (frames / second) 640x480 motion JPEG video requires data at 40Mbps (megabits / second) or higher. Such high data rates for such low resolution video windows are exorbitantly expensive in many broadband applications (and certainly for most consumer Internet-based applications). Moreover, since each frame is compressed independently, defects within the frame that can result from lossy compression will probably appear at different locations in successive frames. This is visible to the viewer as a moving visual flaw when the video is unzipped.
0062Other compression algorithms such as MPEG2, H.264, or VC9 from Microsoft can achieve high compression ratios when used in traditional configurations, but at the expense of long latency. Such algorithms use interframe and intraframe compression. Periodically, such an algorithm performs intraframe-only compression of frames. Such frames are known as keyframes (typically referred to as "I" frames). These algorithms then typically compare I-frames with both front and consecutive frames. Rather than compressing the foreground and contiguous frames independently, this algorithm determines what has changed in the video from the I frame to the foreground and contiguous frames, and then changes those changes. Changes that precede the I frame are stored as "B" frames, and changes that follow the I frame are stored as "P" frames. This results in a much slower data rate than intraframe-only compression. However, this typically comes at the expense of long wait times. I-frames are typically much larger (often 10x or more) than B or P frames, resulting in a proportionally longer time to transmit at a given data rate.
0063For example, the I frame is 10 times the size of the B and P frames, and there are 29 B frames + 30 P frames = 59 interframes per single I intraframe, or "frames". Consider a situation where there are a total of 60 frames for each "group (GOP)". So at 60fps, there is one 60-frame GOP per second. Assume that the transmit channel has a maximum data rate of 2 Mbps. To get the highest quality video on the channel, the compression algorithm produces a 2Mbps data stream, which, assuming the ratios mentioned above, is 2 megabits (Mb) / (59 + 10) = 30,394 bits / intra. Generate frames and 303,935 bit / I frames. When the compressed video stream is received by the decompression algorithm, each frame must be decompressed and displayed at regular intervals (eg 60fps) in order for the video to play steadily. To obtain this result, if a frame is subject to transmission latency, then all frames must be delayed by at least that latency, so the worst-case frame latency is the latency for each video frame. To define. Since the I-frame is the largest, it must introduce the longest transmit latency and receive the entire I-frame before it can be decompressed and displayed (or the interframe depends on the I-frame). .. Assuming a channel data rate of 2 Mbps, it would take 303,935 / 2Mb = 145ms to transmit an I-frame.
0064The above-mentioned interframe video compression system, which uses most of the bandwidth of the transmission channel, is subject to long latency due to the large size of the I-frame relative to the average size of the frame. Or, in other words, traditional interframe compression algorithms achieve lower average data rates per frame than intraframe-only compression algorithms (eg, 2 Mbps vs. 40 Mbps), but still higher per frame due to the large I-frames. I am suffering from peak data rates (eg 303,935 * 60 = 18.2Mbps). Note that the analysis assumes that both the P and B frames are much smaller than the I frame. This is generally true, but not true for frames in which high visual complexity does not correlate with the foreground, high motion, or scene changes. In such a situation, the P or B frame is as large as the I frame (when the P or B frame is larger than the I frame, a sophisticated compression algorithm typically "forces" the I frame. ) , And replace the P or B frame with an I frame). Therefore, the digital video stream always has an I-frame size data rate peak. Thus, for compressed video, when the average video data rate approaches the data rate capacity of the transmit channel (which is often the case given high data rate requirements for video), from an I frame or a large P or B frame. High peak data rates cause long frame latency.
0065Of course, the above description only characterizes the latency of the compression algorithm caused by the large B, P or I frames in the GOP. If B frames are used, the latency will be longer. The reason is that the B frame and all B frames after the I frame must be received before the B frame can be displayed. Therefore, in an image group (GOP) sequence such as BBBBBIPPPPPBBBBBIPPPPP with 5 B frames before each I frame, the first B frame is displayed by the video decompressor until subsequent B and I frames are received. Can not do it. So if the video is streamed at 60fps (ie 16.67ms / frame), then no matter how fast the channel bandwidth is, the 5 B and I frames will be before the first B frame can be decompressed. It takes 16.67 * 6 = 100ms to receive, which is exactly 5 B frames. Compressed video sequences with 30 B-frames are quite common. And at low channel bandwidths such as 2 Mbps, the effect of latency caused by the size of the I frame is primarily added to the effect of latency caused by waiting for the B frame to arrive. Therefore, in a 2 Mbps channel with a very large number of B frames, it is extremely easy to exceed a wait time of 500 ms or more using conventional video compression technology. If B frames are not used (at the expense of a low compression ratio for a given quality level), they do not suffer from B frame latency, but the latency caused by the peak frame size described above remains. suffer.
0066This problem is exacerbated by the very nature of many video games. Video compression algorithms using the GOP structure described above are optimized for use, primarily in live video or video material intended for passive viewing. Typically, the camera (whether it is a real camera or a virtual camera in the case of computer-generated animation), and the scene are relatively stable. This is because the video or movie material is (a) typically uncomfortable to watch if the camera or scene moves too quickly, and (b) the camera suddenly squeaks when watching it. This is because the viewer usually cannot follow the action exactly when moving with (for example, when shooting a child breathing on a birthday cake candle, the camera is bumped and suddenly returns away from the cake. When moved with a gui, the viewer typically concentrates on the child and the cake, ignoring short interruptions when the camera suddenly moves). In the case of video interviews or video teleconferencing, the camera is held in a fixed position and does not move at all, producing very few data peaks. However, 3D high-action video games are characterized by constant movement (for example, consider a 3D race where all frames are in rapid movement during the race, or the virtual camera is constantly slamming. Think of a first-person shooter). Such video games give rise to a frame sequence with large and frequent peaks where the user needs to clearly see what happened during these sudden movements. Therefore, in 3D high-action video games, compression defects are almost unacceptable. Therefore, the video output of many video games, due to their nature, produces compressed video streams with very high and frequent peaks.
0067Assuming that users of high-speed action video games are not very tolerant of long wait times, and given all the causes of the wait times mentioned above, to date, server-hosted streaming video over the Internet. Video games were limited. In addition, users of applications that require a high degree of interactivity are similarly plagued by limitations when the application is hosted on the common Internet and stream video. Such services are such that the hosting server is in a commercial setting, headend (for cable broadband), or central office (for digital subscriber line (DSL)), or LAN (or special grade private connection). It requires a network configuration that is set directly within to control the route and distance from the client device to the server to minimize latency and to accept peaks without latency. A LAN (typically rated from 100 Mbps to 1 Gbps) and a rental line with sufficient bandwidth can typically support peak bandwidth requirements (eg, a peak bandwidth of 18 Mbps is a small LAN capacity of 100 Mbps). Part).
0068Also, peak bandwidth requirements can be accepted by residential broadband infrastructure if special acceptance is made. For example, in a cable TV system, digital video traffic is given a dedicated bandwidth that can handle peaks such as large I-frames. And in DSL systems, high-speed DSL modems can be prepared to allow high peaks or special grade connections that can handle high data rates. However, traditional cable modem and DSL infrastructures installed on the general Internet have much less tolerance for the peak bandwidth requirements of compressed video. Therefore, online services that host video games or applications in a server center long distance from the client device and then stream the compressed video output over the Internet over a traditional residential broadband connection have very short latency, in particular. Demanding games and applications (eg, first person shooters and other multi-user interactive action games, or applications that require fast response times) are plagued by significant latency and peak bandwidth limitations.
0069This disclosure will be more fully understood from the accompanying drawings and the detailed description below. However, the gist disclosed herein is merely an example of the present invention and is not limited to the specific embodiments shown herein.
0070<figref num="1">Shows the architecture of a traditional video game system.</figref><figref num="2a">A high level system architecture according to one embodiment is shown.</figref><figref num="2b">A high level system architecture according to one embodiment is shown.</figref><figref num="3">Shows the actual data rate, rated data rate, and required data rate for communication between the client and server.</figref><figref num="4a">Shows the hosting service and client used by one embodiment.</figref><figref num="4b">Shown is an exemplary latency associated with communication between a client and a hosting service.</figref><figref num="4c">A client device according to an embodiment is shown.</figref><figref num="4d">A client device according to another embodiment is shown.</figref><figref num="4e">It is a block diagram of the client device of FIG. 4c.</figref><figref num="4f">It is a block diagram of the client device of FIG. 4d.</figref><figref num="5">An example is an example of a form of video compression that can be used by an embodiment.</figref><figref num="6a">Illustrates a form of video compression that can be used in another embodiment.</figref><figref num="6b">Shows peak data rates associated with sending low complexity, low action video sequences.</figref><figref num="6c">Shows peak data rates associated with sending high complexity, high action video sequences.</figref><figref num="7a">The video compression technique used in one embodiment is shown.</figref><figref num="7b">The video compression technique used in one embodiment is shown.</figref><figref num="8">An additional video compression technique used in one embodiment is shown.</figref><figref num="9a">The technique used in one embodiment to mitigate the data rate peak is shown.</figref><figref num="9b">The technique used in one embodiment to mitigate the data rate peak is shown.</figref><figref num="9c">The technique used in one embodiment to mitigate the data rate peak is shown.</figref><figref num="10a">An embodiment of efficiently packing video tiles in a packet is shown.</figref><figref num="10b">An embodiment of efficiently packing video tiles in a packet is shown.</figref><figref num="11a">An embodiment using a forward error correction technique is shown.</figref><figref num="11b">An embodiment using a forward error correction technique is shown.</figref><figref num="11c">An embodiment using a forward error correction technique is shown.</figref><figref num="11d">An embodiment using a forward error correction technique is shown.</figref><figref num="12">An embodiment showing a multi-core processing unit for compression is shown.</figref><figref num="13a">Geographical positioning according to embodiments and communication between hosting services are shown.</figref><figref num="13b">Geographical positioning according to embodiments and communication between hosting services are shown.</figref><figref num="14">Illustrate the latency associated with communication between a client and a hosting service.</figref><figref num="15">Shows the server center architecture of the hosting service.</figref><figref num="16">Illustrates a screenshot of an embodiment of a user interface that includes multiple raw video windows.</figref><figref num="17">The user interface of FIG. 16 after selecting a specific video window is shown.</figref><figref num="18">The user interface of Figure 17 after zooming a particular video window to full screen size is shown.</figref><figref num="19">Illustrates collaborative user video data overlaid on the screen of a multiplayer game.</figref><figref num="20">Illustrate a user page for a game player in a hosting service.</figref><figref num="21">Illustrate 3D two-way advertising.</figref><figref num="22">Illustrates a series of steps for generating a photorealistic image with a textured surface from a raw acting surface capture.</figref><figref num="23">Illustrate a user interface page that allows selection of linear media content.</figref><figref num="24">It is a graph which shows the length of time elapsed before a web page is raw with respect to the connection speed.</figref>
0071In the following description, specific details such as device type, system configuration, communication method, etc. are described in order to fully understand the present disclosure. However, it will be apparent to those skilled in the art that these particular details are not required to embody the embodiments described herein.
0072Figure 2a-b shows the video game and software application via the Internet 206 (or other public or private network) under contract service to the user's house 211 (user's house is where the user is located). Indicates a high-level architecture of two embodiments, such as hosted by hosting service 210 and accessed by client device 205 in mobile devices (including outdoors when used). Client device 205 may be a general purpose computer such as a Microsoft Windows or Linux-based PC or Apple's Macintosh computer that has an internal or external display device 222 and is wired or wirelessly connected to the Internet, or video and audio. It may be a dedicated client device such as a set-top box (wired or wirelessly connected to the Internet) that outputs the device to a monitor or TV receiver 222, or perhaps a mobile device that is wirelessly connected to the Internet.
0073Each of these devices may have its own user input device (eg, keyboard, button, touch screen, trackpad or inertial sensing rod, video capture camera and / or motion tracking camera, etc.). , Or a wired or wirelessly connected external input device 221 (eg, keyboard, mouse, game controller, inertial sensing rod, video capture camera, and / or motion tracking camera, etc.) may be used. As detailed below, the hosting service 210 includes servers of various performance levels, including those with high power CPU / GPU processing power. While playing a game or using an application on the hosting service 210, the home or office client device 205 receives a keyboard and / or controller input from the user and then sends the controller input to the hosting service 210 through the internet 206. Sending, this hosting service 210 executes the game program in response, and generates a series of frames (a series of video footage) of video output for the game or application software (eg, the user presses a button, When instructing a character on the screen to move to the right, the game program produces a series of video footage showing the character moving to the right). This series of video footage is then compressed using a short latency video compressor, and then the hosting service 210 sends a short latency video stream over the Internet 206. The home or office client device then decodes the compressed video stream and renders the decompressed video footage on a monitor or TV. As a result, the computing and graphics hardware requirements of client device 205 are significantly relaxed. Client 205 forwards keyboard / controller input to Internet 206 and receives compression from Internet 206. It only needs to have the processing power to decode and decompress the video stream, which is effectively what a personal computer can do today on its CPU with soft care (eg, Intel running at nearly 2GHz). The company's Core Duo CPU can decompress encoded 720p HDTV using compressors like H.264 and Windows Media VC9). And in the case of client devices, dedicated chips also decompress videos to such standards in real time, at a much lower cost, and consume far less power than general-purpose CPUs such as those required for modern PCs. Can be carried out with. In particular, to perform the function of transferring controller input and decompressing video, the home client device 205 is a specialized graphics processing unit (GPU), optical drive or hard drive, eg, a conventional conventional one shown in FIG. Does not require a video game system. You can decompress an encoded 720p HDTV using a compressor like 264 and Windows Media VC9). And in the case of client devices, dedicated chips also decompress videos to such standards in real time, at a much lower cost, and consume far less power than general-purpose CPUs such as those required for modern PCs. Can be carried out with. In particular, to perform the function of transferring controller input and decompressing video, the home client device 205 is a specialized graphics processing unit (GPU), optical drive or hard drive, eg, a conventional conventional one shown in FIG. Does not require a video game system. You can decompress an encoded 720p HDTV using a compressor like 264 and Windows Media VC9). And in the case of client devices, dedicated chips also decompress videos to such standards in real time, at a much lower cost, and consume far less power than general-purpose CPUs such as those required for modern PCs. Can be carried out with. In particular, to perform the function of transferring controller input and decompressing video, the home client device 205 is a specialized graphics processing unit (GPU), optical drive or hard drive, eg, a conventional conventional one shown in FIG. Does not require a video game system.
0074As game and application software becomes more complex and more holistic, they require high performance CPUs, GPUs, more RAM, and larger and faster disk drives, and hosting services 210. Computing power continues to upgrade, but end users are not required to update their home or office client platform 205. This is because its processing requirements remain constant for the display resolution and frame rate of a given video decompression algorithm. Therefore, the system shown in Figures 2a-b does not have the hardware limitation and compatibility issues found today.
0075Furthermore, since the game and application software runs only on the server of hosting service 210, a copy of the game or application software (in the form of optical media or as downloaded software) may be present in the user's home or office. None ("Office" as used herein includes non-residential settings, including, for example, school classrooms, unless otherwise noted). This significantly reduces the risk of illegally copying (pirated) game or application software, as well as reducing the risk of valuable databases being used by pirated games or applications. In fact, when special servers are required to play games or application software that are impractical for home or office use (eg, very expensive, loud or noisy equipment is required). Even if a pirated copy of the game or application software is obtained, it cannot operate at home or in the office.
0076In one embodiment, the hosting service 210 provides software development tools to a game or application software developer (generally referring to a software developer, game or movie studio, or game or application software publisher) 220 and this development. A person designs a video game to design a game that can be run on the hosting service 210. Such tools make hosting service features that are not normally available on stand-alone PCs or video game consoles available to developers (eg, fast access to very large databases with complex geometry (geometry). "Geometry" here refers to polygons, textures, rigging, lighting, behavior, and other components and parameters that define a 3D database, unless otherwise noted)).
0077Under this architecture, different business models are possible. Under one model, the hosting service 210 collects contract fees from end users and pays royalties to developers 220, as shown in Figure 2a. In another embodiment shown in Figure 2b, the developer 220 collects the contract fee directly from the user and pays the hosting service 210 to host the game or application content. These basic principles are not limited to a particular business model for providing online gaming or application hosting.
0078<u style="single">Compressed video characteristics</u> As mentioned above, one notable problem with providing video game services or application software services online is latency. A wait time of 70 to 80 ms (from the point where the input device is operated by the user to the point where the response is displayed on the display device) is the upper limit for games and applications that require fast response times. However, this is very difficult to achieve due to the large number of practical and physical constraints in the architectural environment shown in Figures 2a and 2b.
0079As shown in FIG. 3, when a user subscribes to an Internet service, the connection is typically rated by a nominal maximum data rate of 301 to the user's home or office. Based on the provider's policies and the capabilities of the routing device, its maximum data rate can be enforced somewhat rigorously, but typically the actual data rate obtained will be low, for one of many different reasons. .. For example, there may be too much network traffic in the DSL central office or local cable modem loop, or the cable may be noisy and drop packets, or the provider may establish a maximum number of bits / month / user. .. Currently, maximum downstream data rates for cable and DSL services typically range from hundreds of kilobits per second (Kbps) to 30 Mbps. Cellular services are typically limited to hundreds of Kbps of downstream data. However, the speed of broadband services and the number of users who subscribe to broadband services increase rapidly over time. Currently, one analysis estimates that 33% of US broadband subscribers have downstream data rates of 2 Mbps and above. For example, one analysis estimates that by 2010, more than 85% of US broadband subscribers will have a data rate of 2 Mbps or higher.
0080As shown in FIG. 3, the maximum data rate 302 actually available can fluctuate over time. Therefore, in short latency online game or application software contexts, it is sometimes difficult to predict the actually available data rate for a particular video stream. A data rate of 303 is required to maintain a given level of quality at a given number of frames per second (fps) at a given resolution (eg, 640x480 @ 60fps) for a certain amount of scene complexity. , And if the movement rises above the maximum data rate 302 that is actually available (as shown by the peak in Figure 3), a number of problems can occur. For example, some internet services simply drop packets, causing data loss and video distortion / loss on the user's video screen. Other services temporarily buffer (ie, queue) additional packets and feed the packets to the client at the available data rate, increasing latency, ie for many video games and applications. It will lead to unacceptable results. Ultimately, some Internet service providers consider increasing data rates to be malicious attacks, such as denial of service attacks (a well-known technique used by hackers to disable network connections), and Disconnect the user's Internet connection for a specified period of time. Therefore, the embodiments described herein take steps to ensure that the data rate required for the video game does not exceed the maximum available data rate.
0081<u style="single">Hosting service architecture</u> FIG. 4a shows the architecture of the hosting service 210 according to one embodiment. The hosting service 210 can be located in a single server center or distributed across multiple server centers (short latency connections where the route to one server center has a shorter latency than other server centers). To give users, load balance between users, and redundancy if one or more server centers fail). The hosting service 210 can ultimately serve a very large user base, including hundreds or thousands or millions of servers 402. The hosting service control system 401 gives overall control over the hosting service 210 and directs routers, servers, video compression systems, billing and accounting systems, and so on. In one embodiment, the hosting service control system 401 is embodied in a distributed processing Linux-based system coupled to a RAID array used to store a database for user information, server information and system statistics. In the above description, the various actions embodied by the hosting service 210 are initiated and controlled by the hosting service control system 401 unless it is caused by another specific system.
0082Hosting service 210 includes a number of servers 402, such as those currently available from Intel, IBM, Hewlett-Packard, etc. Alternatively, the server 402 can be assembled with a custom configuration of components, or it can be finally integrated so that all servers are embodied as a single chip. This figure shows only a small number of servers 402 for illustration clarity, but in a real deployment there may be as many as one server 402, or millions or more of servers 402. May be good. Server 402 may all be configured in the same way (as an example of some configuration parameters, with the same CPU type and performance, with or without a GPU, and if it has a GPU, With the same GPU format and performance, with the same number of CPUs and GPUs, with the same amount and format / speed of RAM, and with the same RAM configuration), or various subsets of server 402 have the same configuration. May (for example, 25% of the servers can be configured in one way, 50% can be configured differently, and 25% can be configured in yet another way), or each The server 402 may be different.
0083In one embodiment, the server 402 is diskless, i.e. its own local mass storage device (optical or magnetic storage device, or semiconductor-based storage device, such as flash memory or other mass storage means performing similar functions. ), Each server accesses shared mass storage through a high-speed backplane or network connection. In one embodiment, this high speed connection is a storage area network (SAN) 403 connected to a series of independent disk redundant arrays (RAID) 405, and the connections between the devices are implemented using Gigabit Ethernet®. Be made. As will be apparent to those skilled in the art, the SAN403 synthesizes a large number of RAID arrays 405 together to eventually produce a large bandwidth, i.e. the bandwidth obtained from the RAM used in current game consoles and PCs. Used to approach or potentially exceed. And while RAID arrays based on rotating media such as magnetic media often have significant seek-time access latency, RAID arrays based on semiconductor storage devices can be embodied with fairly short access latency. Can be done. In another configuration, some or all of the servers 402 provide some or all of their own mass storage locally. For example, server 402 stores frequently accessed information, such as its operating system, and a copy of a video game or application in a short-latency local flash-based storage device, but for geometric shape or game state information. To access large databases from time to time, use the SAN to access the rotating media-based RAID array 405 with high seek latency.
0084Further, in one embodiment, the hosting service 210 uses the short latency video compression logic 404 described in detail below. This video compression logic 404 can be embodied in software, hardware or a combination thereof (some embodiments thereof will be described below). Video compression logic 404 includes logic for compressing audio and visual material.
0085In operation, the control signal logic 413 of the client 415 is activated by the user while playing a video game or using an application in the user's house 211 via a keyboard, mouse, game controller or other input device 421. It sends a control signal 406a-b (typically in the form of a UDP packet) that represents a button press (and other forms of user input). Control signals from a given user are routed to the appropriate server (or multiple servers if multiple servers respond to the user's input device) 402. As shown in FIG. 4a, the control signal 406a is routed through the SAN to the server 402. Separately or additionally, the control signal 406b is routed directly to the server 402 via a hosting service network (eg, an Ethernet-based local area network). Regardless of how they are transmitted, the server (s) runs the game or application software in response to the control signal 406a-b. Although not shown in Figure 4a, various network components such as firewalls (s) and / or gateways (s) are associated with hosting service 210 (eg, hosting service 210 and internet 410). In (between) and / or at the edge of the user's home 211 between the Internet 410 and the home or office client 415, it can handle incoming and outgoing traffic. The graphic and audio output of the executed game or application software, a new sequence of video footage, is fed to the short latency video compression logic 404, which logic is a short latency video compression technique as described herein. A sequence of video footage is compressed based on, and the compressed video stream is typically compressed or uncompressed. -Return to client 415 via the Internet 410 (or via optimal high-speed network services that bypass the general Internet, as described below) with the Dio. The short latency video decompression logic 412 on the client 415 then decompresses the video and audio streams, renders the decompressed video stream, and typically plays the decompressed audio stream on display device 422. To do. Alternatively, the audio may or may not be played on speakers separate from the display device 422. Note that the input device 421 and the display device 422 are shown as independent devices in FIGS. 2a and 2b, but may be integrated within a client device such as a portable computer or mobile device.
0086The home or office client 415 (already mentioned as the home or office client 205 in Figures 2a and 2b) is a very inexpensive and low power device with very limited computational or graphic performance and a large number of locals. The storage device is very limited or does not have it at all. In contrast, SAN403 and each server 402 coupled to multiple RAID405s is a very high performance computing system, and in fact if multiple servers are used collaboratively in a parallel processing configuration. There is almost no limit to the amount of computing and graphics processing power that can be retained. The computational power of the server 402 is then given to the user for the short latency video compression 404 and the short latency video compression 412 perceptually to the user. When the user presses a button on input device 421, the image on display 422 responds to the button press with no perceptually significant delay, as if the game or application software were running locally. Will be updated. Therefore, in a home or office client 415, which is a very low performance computer or inexpensive chip that embodies short latency video decompression and control signal logic 413, any effective remote location that may be available locally. Computational power is given to the user. This gives users the power to play the most advanced processor-intensive (typically new) video games and top-performing applications.
0087FIG. 4c shows a very basic and inexpensive home or office client device 465. This device is an embodiment of home or office client 415 from FIGS. 4a and 4b. This is about 2 inches long. It has an Ethernet jack 462 that interfaces with an Ethernet cable over Power over Ethernet (PoE), from which power and connectivity to the Internet can be obtained. You can perform NAT within a network that supports Network Address Translation (NAT). In an office environment, many new Ethernet switches have PoE and bring PoE directly to the office Ethernet jack. In such a situation, only an Ethernet cable from the wall jack to the client 465 is required. If the available Ethernet connection does not carry power (eg, in a home with a DSL or cable modem but no PoE), a cheap wall "brick" that accepts unpowered Ethernet cables and output Ethernet with PoE. (Ie power supply) is available.
0088Client 465 includes a control signal logic 413 (Figure 4a) coupled to a Bluetooth wireless interface that interfaces with a Bluetooth input device 479 such as a keyboard, mouse, game controller and / or microphone and / or headset. Also, one embodiment of the client 465 outputs video coupled to a display device 468 capable of supporting 120 fps video at 120 fps and signals the shuttered eyeglasses 466 (typically by infrared), one after the other. The frame allows the shutter to be activated alternately in one eye and then in the other. The effect that the user perceives is a stereoscopic 3D image that "jumps out" the display screen. One such display device 468 that supports such operations is Samsung It is HL-T5076S. Since each eye video stream is separate, in one embodiment two independent video streams are compressed by the hosting service 210, frames are interleaved in time, and frames are independent 2 in client 465. It is decompressed as one decompression process.
0089The client 465 also has a short latency video decompression logic 412 that decompresses incoming video and audio and outputs it through HDMI (High Definition Multimedia Interface), and SDTV (Standard Sharpness Television) or HDTV (High Definition Television) or HDTV (High Definition Television). Degree Television) Includes a connector 463 that plugs into a 468 to provide video and audio to the TV or plugs into a monitor 468 that supports HDMI. If the user's monitor 468 does not support HDMI, HDMI vs. DVI (Digital Visual Interface) can be used, but audio is lost. Under the HDMI standard, display capabilities (eg, supported resolutions, frame rates) 464 are communicated from display device 468, and this information is then sent back to the hosting service 210 over the internet connection 462, thus displaying. Compressed video can be streamed in a format suitable for the device.
0090FIG. 4d shows a home or office client device 475 that is the same as the home or office client device 465 shown in FIG. 4c except that it has more external interfaces. The client 475 can also accept PoE for power or extend from an external power adapter (not shown) that plugs into the wall. Using the client 475's USB input, the camcorder 477 feeds the compressed video to the client 475, which is uploaded by the client 475 to the hosting service 210 and used as described below. Built into the camera 477 is a short-latency compressor that uses the compression techniques described below.
0091In addition to having an Ethernet connector as an internet connection, the client 475 also has an 802.11g wireless interface to the internet. Both interfaces can use NAT in networks that support NAT.
0092In addition to having an HDMI connector for outputting video and audio, the client 475 also has a dual link DVI-I connector that includes an analog output (and provides VGA output with a standard adapter cable). It also has analog outputs for composite video and S-video.
0093For audio, the client 475 has left and right analog stereo RCA jacks, and for digital audio output, a TOSLINK output.
0094It also has a USB jack for interfacing the input device, in addition to the Bluetooth wireless interface to the input device 479.
0095FIG. 4e shows an embodiment of the client 465's internal architecture. All or some of the illustrated devices can be embodied in field programmable logic arrays, custom ASICs, or a large number of custom designed or off-the-shelf individual devices.
0096Ethernet 497 with PoE is installed on Ethernet interface 481. Power 499 is derived from Ethernet 497 with PoE and connected to the rest of the equipment in client 465. Bus 480 is a common bus for communication between devices.
0097Control CPU 483 running a small client control application from flash 476 (most often 100MHz MIPS with embedded RAM A small CPU, such as the R4000 series CPU, will suffice) to embody the protocol stack for the network (ie, the Ethernet interface), also communicate with the hosting service 210, and all the equipment in the client 465. Configure. It also handles the interface with the input device 469 and, if necessary, protects it with "forward error correction" and sends the packet back to the hosting service 210 along with the user controller data. The control CPU 483 also monitors packet traffic (for example, if packets are lost or delayed, they also time stamp their arrival). This information is sent back to the hosting service 210, so you can constantly monitor your network connection and adjust what you send accordingly. The flash memory 476 is initially loaded with the control program of the control CPU 483 at the time of manufacture, as well as a serial number unique to a particular client 465 unit. This serial number allows the hosting service 210 to uniquely identify the client 465 unit.
0098Bluetooth interface 484 wirelessly communicates with input device 469 through its antenna inside client 465.
0099The video decompressor 486 is a short latency video decompressor configured to embody the video decompression described herein. A large number of video decompression devices exist as off-the-shelf or as intellectual property (IP) designs that can be integrated into FPGAs or custom ASICs. One company that provides IP for H.264 decoders is NSW Australia's Ocean Logic of Manly. The effect of using IP is that the compression technique used here does not comply with compression standards. Some standard decompressors are flexible enough to be configured to accept the compression techniques described here, while others are not. However, IP has complete flexibility in redesigning the decompressor as needed.
0100The output of the video decompressor is coupled to the video output subsystem 487, which couples the video to the video output of HDMI interface 490.
0101The audio decompression subsystem 488 can be embodied using standard audio decompressors available, or can also be embodied as an IP, or audio decompression embodies, for example, the Vorbis audio decompressor. It can also be embodied within the control processor 483.
0102The device that embodies audio decompression is coupled to the audio output subsystem 489, which couples the audio to the audio output of HDMI interface 490.
0103Figure 4f shows an embodiment of the client 475's internal architecture. Obviously, this architecture is the same as the client 465, except for the additional interface and any external DC power from the wall-plugged power adapter, and that external DC power is used as such. If so, it will replace the power coming from the Ethernet PoE497. The functions common to the client 465 will not be repeated below, but additional functions will be described below.
0104CPU483 communicates with and configures additional equipment.
0105WiFi subsystem 482 provides wireless internet access as an alternative to Ethernet 497 through its antenna. WiFi subsystems are available from a wide range of manufacturers, including Atheros Communications in Santa Clara, California.
0106The USB subsystem 485 provides an alternative to Bluetooth communication for the wired USB input device 479. USB subsystems are quite standard, readily available for FPGAs and ASICs, and are often built into off-the-shelf equipment that performs other functions as well as video decompression.
0107The video output subsystem 487 produces a wider range of video output than within the client 465. This provides DVI-I491, S-Video 492 and Composite Video 493 in addition to providing HDMI 490 video output. Also, when the DVI-I491 interface is used for digital video, the display capability 464 is returned from the display device to the control CPU 483 so that the capability of the display device 478 can be notified to the hosting service 210. All interfaces provided by the video output subsystem 487 are very standard interfaces and are readily available in many forms.
0108The audio output subsystem 489 digitally outputs audio through the digital interface 494 (S / PDIF and / or Toslink) and outputs audio in analog form through the stereo analog interface 495.
0109<u style="single">Round-trip waiting time analysis</u> Of course, in order to understand the above paragraph, the round-trip wait time between the user's action using input device 421 and viewing the result of that action on display device 420 must be 70-80ms or less. .. This latency must take into account all factors in the path from the input device 421 in the user's house 211 to the hosting service 210 and back to the user's house 211 to the display device 422. Figure 4b shows the various components and networks in which the signal must travel, and above these components and networks is a timeline listing the latency that can be expected in the actual realization. Note that Figure 4b has been simplified to show only important route routing. Other routing of data used for other features of the system is described below. A double-headed arrow (eg, arrow 453) represents a round-trip wait time, a single-headed arrow (eg, arrow 457) represents a one-way wait time, and "~" represents an approximate measure. Although there are real-world situations where the listed latency cannot be achieved, in many cases in the United States, DSL and cable modem connections to the user's home 211 are used to achieve these latency in the environment described in the next paragraph. We must point out what we can do. Also, while cellular wireless connections to the Internet work reliably in the systems shown, most current US cellular data systems (such as EVDO) suffer very long wait times, as shown in Figure 4b. Also note that time cannot be achieved. However, these basic principles could be embodied in future cellular technologies that can embody this level of latency.
0110Starting from the input device 421 in the user's house 211, when the user operates the input device 421, a user control signal is sent to the client 415 (which may be a stand-alone device such as a set-top box, or a PC. Or software or hardware running on another device, such as a mobile device), and packetized (in UDP format in one embodiment), the packet being the destination address for arriving at the hosting service 210. Is given. The packet also includes information indicating from which user the control signal comes from. The control signal packet (s) are then forwarded through the firewall / router / NAT (Network Address Translation) device 443 to WAN interface 442. WAN interface 442 is an interface device provided to the user's house 211 by the user's ISP (Internet Service Provider). The WAN interface 442 may be a cable or DSL modem, a WiMax transceiver, a fiber transceiver, a cellular data interface, an Internet Protocol over powerline interface, or many other interfaces to the Internet. In addition, the firewall / router / NAT device 443 (and potentially WAN interface 442) may be integrated with the client 415. One example is a mobile phone that includes software for embodying the functionality of a home or office client 415 and means for wirelessly routing and connecting to the Internet according to certain standards (eg 802.11g).
0111WAN interface 442 then sends the control signal to the user's Internet Service Provider (ISP) at the "point of point of point of Routed to what is referred to as "presence)" 441, which is the facility that interfaces between the WAN transport connected to the user's home 211 and the general Internet or private network. The characteristics of the point of existence vary based on the nature of the Internet services provided. In the case of a DSL, this is typically the central office of the telephone company where the DSLAM is located. For cable modems, this is typically a cable multisystem operator (MSO) headend. For cellular systems, this is typically the control room associated with the cellular tower. However, whatever the nature of the point of existence, it routes control signal packets (s) to the general Internet 410. The control signal packets (s) are then routed to WAN interface 441 to hosting service 210, most often through fiber transceiver interfaces. WAN441 then routes the control signal packet to routing logic 409, which is embodied in many different ways, including Ethernet switches and routing servers, which evaluates the user's address and then Route the control signal to the correct server 402 for a given user.
0112The server 402 then takes the control signal as input to the game or application software running on the server 402 and uses the control signal to process the next frame of the game or application. When the next frame occurs, video and audio are output from the server 402 to the video compressor 404. The video and audio are output from the server 402 to the compressor 404 via various means. First, the compressor 404 can be incorporated into the server 402, so compression can be embodied locally within the server 402. Alternatively, video and / or audio, in packet form, via a network connection such as an Ethernet connection, to a network that is a private network between the server 402 and the video compressor 404, or through a shared network such as SAN403. Can be output. Alternatively, the video may be output from the server 402 through a video output connector such as a DVI or VGA connector and then captured by the video compressor 404. The audio may also be output from the server 402 as digital audio (eg, via a TOSLINK or S / PDIF connector) or as analog audio, which is produced by the audio compression logic in the video compressor 404. It is digitized and encoded.
0113When the video compressor 404 captures a video frame and the audio generated during that frame time from the server 402, it compresses the video and audio using the techniques described below. When the video and audio are compressed, they are packetized with the address and sent back to the user's client 415 and routed to WAN interface 441, which routes video and audio packets through the common Internet 410. The Internet routes video and audio packets to the user's ISP location 441, which routes video and audio packets to the WAN interface 442 of the user's home, which interfaces the video and audio packets. It routes to firewall / router / NAT device 443, which then routes video and audio packets to client 415.
0114Client 415 decompresses the video and audio, then displays the video on display device 422 (or the client's built-in display device) and displays the audio to display device 422 or a separate amplifier / speaker or to the client's built-in amplifier. / Send to speaker.
0115The round-trip delay must be less than 70 or 80 ms in order for the user to perceive that there is no perceptual delay in all of the processes described above. Those with latency delays on the round-trip route are under the control of hosting service 210 and / or the user, others are not. Nevertheless, based on the analysis and testing of a large number of real-world scenarios, the approximate measurements are:
0116The one-way transmission time to transmit the control signal 451 is typically less than 1ms, and the round-trip routing through the user's house 452 is typically a consumer-grade firewall / router / readily available over Ethernet. Achieved in about 1ms using a NAT switch. User ISPs vary widely in their round-trip delay 453, but DSL and cable modem providers typically have 10 to 25 ms. Round-trip latency on a typical Internet 410 varies widely based on how traffic is routed and whether the route is flawed (these issues are described below), but is typically common. The Internet gives a very optimal route, and latency is determined primarily by the speed of light through fiber optics, given the distance to the destination. As further described below, we have established 1000 miles as the farthest approximate distance expected to launch hosting service 210 away from the user's home 211. At 1000 miles (2000 miles round trip), the actual transit time of a signal through the Internet is about 22 ms. WAN interface 441 to hosting service 210 is typically a high speed interface of commercial grade fiber with negligible latency. Therefore, a typical internet latency 454 is typically 1 to 10 ms. The latency of one-way routing 455 through hosting service 210 is less than 1ms. Server 402 typically puts a new frame for a game or application in one frame time (16 at 60 fps). Calculated in less than 7ms), so the maximum reasonable one-way latency to use is 16ms. Optimal hardware implementation of the video compression and audio compression algorithms described here can complete compression 457 in 1 ms. In the less optimal form, compression takes about 6ms (of course, the less optimal form takes longer, but such realization affects the total round-trip waiting time, waiting 70-80ms. Other wait times need to be reduced to maintain the time goal (for example, the permissible distance through the general Internet can be reduced). The round-trip latency of Internet 454, user ISP453 and user house routing 452 has already been taken into account, so the rest is the latency of video decompression 458, which is embodied by video decompression 458 with dedicated hardware. It varies based on whether it is done or embodied in software in a client device 415 (such as a PC or mobile device), and based on the size of the display and the performance of the decompressed CPU. Defrosting 458 typically takes 1-8ms.
0117Therefore, by adding up all the worst-case waiting times that are actually seen, the worst-case round-trip waiting time that can be expected to be experienced by the users of the system shown in FIG. 4a can be determined. They are 1 + 1 + 25 + 22 + 1 + 16 + 6 + 8 = 80ms. And, indeed (with the notes described below), it uses a prototype version of the system shown in Figure 4a, using an off-the-shelf Windows PC as a client device and a home DSL and cable modem in the United States. Approximate round-trip latency found using the connection. Of course, in better scenarios than in the worst case, very short latency is obtained, but it cannot rely on the development of widely used commercial services.
0118To obtain the latency listed in Figure 4b over the general Internet, the video compressor 404 in Figure 4a and the video decompressor 412 in Client 415 generate a packet stream with very specific characteristics and the hosting service 210. Packet sequences generated over the entire route from to display device 422 are not subject to delays or excessive packet loss, and are especially available to users via the user's internet connection through WAN interface 442 and firewall / router / NAT443. You need to be consistent within the limits of the bandwidth you can afford. In addition, the video compressor must generate a packet stream that is robust enough to tolerate the inevitable packet loss and packet reordering that occurs in normal Internet and network transmission.
0119<u style="single">Short latency video compression</u> To achieve the above goal, one embodiment takes a novel solution of video compression that reduces latency and relaxes the peak bandwidth requirement for transmitting video. Before explaining these embodiments, an analysis of current video compression techniques is performed with reference to FIGS. 5 and 6a-b. Of course, these techniques can be used on the basis of basic principles if the user is given sufficient bandwidth to handle the data rates required by these techniques. Note that audio compression is not covered here except to state that it is embodied simultaneously and synchronously with video compression. There are conventional audio compression techniques that meet the requirements of this system.
0120Figure 5 shows one particular prior art for compressing video, where each individual video frame 501-503 is compressed by a compression logic 520 that uses a particular compression algorithm to form a series of compressed frames. Shows the prior art of generating 511-513. One embodiment of this technique is "Motion JPEG", where each frame is compressed according to a Joint Picture Expert Group (JPEG) compression algorithm based on the Discrete Cosine Transform (DCT). .. A variety of different types of compression algorithms may be used, but will still be adapted to these underlying principles (eg, wavelet-based compression algorithms such as JPEG-2000).
0121One problem with this form of compression is that it reduces the data rate of each frame, but does not take advantage of the similarity between successive frames to reduce the data rate of the entire video stream. For example, assuming a frame rate of 640x480x24 bits / pixel = 640 * 480 * 24/8/1024 = 900 kilobytes / frame (KB / frame) for a given quality of video, as shown in Figure 5. Motion JPEG only compresses the stream by 10 factors, producing a 90KB / frame data stream. At 60 frames / sec, this is 90KB * 8 bits * 60 frames / sec = 42. It requires a channel bandwidth of 2 Mbps, which is much wider for almost every home internet connection in the United States today, and very wide for many office internet connections. It becomes. In fact, if you require a constant data stream with such a wide bandwidth and it is only useful for one user in an office LAN environment, it will consume a large percentage of the 100Mbps Ethernet LAN bandwidth and LAN. It will put a heavy burden on the Ethernet switch that supports. Therefore, compression for moving images and videos is inefficient when compared to other compression techniques (as described below). In addition, single-frame compression algorithms such as JPEG and JPEG-2000, which use the Rossie compression algorithm, do not notice compression defects in still video (eg, defects in dense leaves in a scene, which dense leaves are. It may not be visible as a defect because the eye does not know exactly what it should look like). However, when the scene moves, the defects become noticeable because it is visually detected that the defects change from frame to frame even though the defects are in the area of the scene that is not noticed in the still image. As a result, "background noise" that looks similar to the "snow" noise seen during margin analog TV reception is perceived in the frame sequence. Of course, this form of compression can still be used in some of the embodiments described herein, but generally speaking, high data for a given perceptual quality to avoid background noise in the scene. A rate (ie, a low compression ratio) is required.
0122Compression of video streams is more efficient because H.264, or other formats of compression such as Windows Media VC9, MPEG2 and MPEG4, all take advantage of the similarities between consecutive frames. All of these techniques rely on the same general technique for compressing video. Therefore, although the H.264 standard is described, the same general principle applies to various other compression algorithms. A large number of H.264 compressors and decompressors are available, including the x264 open source software library for compressed H.264 and the FFmpeg open source software library for decompressed H.264.
0123Figures 6a and 6b illustrate conventional compression techniques, where a series of uncompressed video frames 501-503, 559-561 are combined with a compression logic 620 to create a series of "I-frames" 611, 671, "P-frames." It is compressed into "612-613" and "B frame" 670. The vertical axis in Figure 6a generally represents the size of each encoded frame (although the frames are not drawn on the proper scale). As mentioned above, video coding using I-frames, B-frames and P-frames is well known to those of skill in the art. Briefly, I-frame 611 is a DCT-based compression of a completely uncompressed frame 501 (similar to the compressed JPEG video described above). The P-frame 612-613 is generally significantly smaller in size than the I-frame 611. This is because it incorporates the advantages of the data in the foreground I-frame or P-frame, that is, it contains data that indicates changes between the foreground I-frame or P-frame. The B frame 670 is similar to the P frame, but the B frame uses a frame in the subsequent reference frame and a potential frame in the preceding reference frame.
0124In the following description, it is assumed that the desired frame rate is 60 frames per second, each I frame is about 160 Kb, the average P and B frames are 16 Kb, and a new I frame is generated every second. With this set of parameters, the average data rate is 160Kb + 16Kb * 59 = 1.1Mbps. This data rate is well within the maximum data rate for many current broadband internet connections to homes and offices. This technique also tends to avoid the problem of background noise from the encoding dedicated to the intraframe. This is because the P and B frames track the differences between the frames, and compression defects do not tend to appear or disappear frame by frame, alleviating the background noise problem.
0125One problem with the format of compression is that the average data rate is relatively low (eg 1.1 Mbps), but a single I-frame takes many frames to transmit. For example, using prior art, a 2.2 Mbps network connection (eg, from Figure 3a to 2.2 Mbps peak available) is typically available to stream video at 1.1 Mbps with an I frame of 160 Kbps every 60 frames. A DSL or cable modem with a maximum data rate of 302) is sufficient. This is achieved by putting a 1 second video in the decompression queue until the video is decompressed. In 1 second, 1.1 Mb of data is transmitted, which is easily accepted by the maximum available data rate of 2.2 Mbps, even assuming that the available data rate drops by as much as 50% periodically. Unfortunately, this traditional solution causes a 1 second wait time for the video because the receiver has a 1 second video buffer. Such delays are sufficient for many traditional applications (eg, playing linear video), but with much longer latency for fast action video games that cannot tolerate latency greater than 70-80ms. is there.
0126Attempts have been made to eliminate the 1-second video buffer, but not enough latency reduction for high-speed action video games. As an example, the use of the B-frame described above requires the reception of all B-frames and I-frames preceding the I-frame. Assuming that 59 non-I-frames are roughly split between the P-frame and the B-frame, there are at least 29 B-frames, and I-frames are received before the B-frame can be displayed. Therefore, regardless of the available bandwidth of the channel, a delay of 29 + 1 = 30 frames each 1/60 second wide, that is, a waiting time of 500 ms is required. Obviously, this is much too long.
0127Therefore, another solution is to eliminate the B frame and use only the I and P frames. (One result is that the data rate increases for a given quality level, but for consistency in this example, each I-frame is 160Kb and the average P-frame is 16Kb in size. Yes, and therefore continue to assume that the data rate is still 1.1 Mbps.) This solution eliminates the inevitable latency introduced by the B frame. This is because the decoding of each P-frame only depends on previously received frames. The problem with this solution is that on narrow-bandwidth channels, which are typical in most homes and many offices, the transmission of an I-frame has a substantial latency because the I-frame is much larger than the average P-frame. Is to increase. This is shown in Figure 6b. The video stream data rate 624 is the maximum available, except for I frames where the peak data rate 623 required for the I frame far exceeds the maximum available data rate 622 (and also the rated maximum data rate 621). Data rate lower than 621. The data rate required by the P-frame is less than the maximum data rate available. Even if the peak of the maximum available data rate of 2.2 Mbps is steadily maintained at that peak rate of 2.2 Mbps, it takes 160Kb / 2.2Mb = 71ms to send an I frame, and the maximum available data rate. 622 is 50% (1. If it drops (1Mbps), it takes 142ms to send an I-frame. Therefore, the wait time for transmitting an I-frame falls somewhere between 71 and 142 ms. This wait time is added to the wait time shown in FIG. 4b, and in the worst case it is added to 70 ms, so that this is because the image appears on the display device 422 from the time the user operates the input device 421. Up to, a total round-trip waiting time of 141 to 222 ms is generated, which is much higher. And if the maximum available data rate drops below 2.2 Mbps, the wait time will increase further.
0128Also note that in general, ISPs are "jammed" at peak data rates of 623, resulting in severe results that far exceed the available data rates of 622. Devices from different ISPs behave differently, but when receiving packets at data rates well above the available data rate 622, the following behavior becomes quite common among DSL and cable modem ISPs. (a) Defer packets by queuing them (introduce latency), (b) drop some or all of the packets, and (c) disable the connection for a period of time (probably for ISPs) Because it is a malicious attack, such as a "denial of service" attack). Therefore, transmitting a packet stream at all data rates with the characteristics shown in FIG. 6b is not a feasible option. Peak 623 may be queued in hosting service 210 and transmitted at a data rate lower than the maximum available data rate, introducing the unacceptable latency described above.
0129In addition, the video stream data rate sequence 624 shown in Figure 6b is a very "submissive" video stream data rate sequence and is expected to result from compressing video from a video sequence that is significantly unchanged and has little motion. A type of data rate sequence that is commonly used in video teleconferencing, where the camera is in a fixed position and moves very little, and objects in the scene, such as the person sitting in a chair and speaking, show little movement. To be the target).
0130The video stream data rate sequence 634 shown in Figure 6c is typical of what is expected to be visible from a video with much more action, such as that occurring in a video or video game or some application software. .. Note that in addition to the I-frame peak 633, there are also P-frame peaks such as 635 and 636 that are extremely large and often exceed the maximum data rate available. These P-frame peaks are not often as large as the I-frame peaks, but are much too large to be carried by the channel at all data rates, and like the I-frame peaks, they are P-frames. Peaks must be transmitted slowly (thus increasing latency).
0131In a broadband channel (eg, a 100 Mbps LAN, or a private connection with a wide bandwidth of 100 Mbps), the network can tolerate large peaks such as I-frame peak 633 or P-frame peak 636, and in principle, A short waiting time can be maintained. However, such networks are often shared among large numbers of users (eg, in office environments), and such "peak" data is especially when network traffic is routed to private shared connections. (For example, from a remote data center to the office), which affects LAN performance. First, note that this example is a relatively low resolution video stream of 640x480 pixels at 60fps. HDTV streams at 1920x1080 at 60fps are easily handled by modern computers and displays, and displays with 2560x1440 resolution at 60fps are becoming more and more available (eg Apple's 30 "displays) and 60fps. In 1920x1080 high action video sequence, using H.264 compression for moderate quality level 4. Requires 5 Mbps. Assuming an I-frame peak that is 10 times the nominal data rate, it produces a peak below 45 Mbps, but there is still a significant P-frame peak. If a large number of users receive a video stream over the same 100 Mbps network (for example, a private network connection between an office and a data center), how the peaks from the large number of users' video streams are aligned, It is easy to see if it overwhelms the bandwidth of the network and potentially overwhelms the bandwidth of the switch backplane that supports users on the network. Even in the case of Gigabit Ethernet networks, the network or network switch can be overwhelmed if enough users align enough peaks at once. And as 2560x1440 resolution video becomes more mediocre, the average video stream data rate will be 9.5 Mbps, probably resulting in a peak data rate of 95 Mbps. Needless to say, the 100Mbps connection between the data center and the office (which is a very fast connection today) is completely sunk by peak traffic from a single user. Therefore, even if LAN and private network connections are more tolerant of peak streaming video, streaming video with high peaks is not desirable and requires special planning and adaptation by the IT department of the office. It will be 5 Mbps, probably resulting in a peak data rate of 95 Mbps. Needless to say, the 100Mbps connection between the data center and the office (which is a very fast connection today) is completely sunk by peak traffic from a single user. Therefore, even if LAN and private network connections are more tolerant of peak streaming video, streaming video with high peaks is not desirable and requires special planning and adaptation by the IT department of the office. It will be 5 Mbps, probably resulting in a peak data rate of 95 Mbps. Needless to say, the 100Mbps connection between the data center and the office (which is a very fast connection today) is completely sunk by peak traffic from a single user. Therefore, even if LAN and private network connections are more tolerant of peak streaming video, streaming video with high peaks is not desirable and requires special planning and adaptation by the IT department of the office.
0132Of course, for standard linear video applications, these things are not a problem. This is because the data rate is "smoothed" at the transmission point, the data in each frame is below the maximum data rate of 622 available, and the client buffer decompresses the sequence of I, P and B frames. This is because it will be remembered until it is done. Therefore, the data rate across the network is kept close to the average data rate of the video stream. Unfortunately, this is unacceptable for short latency applications such as video games and applications that introduce latency even when B-frames are not used, i.e. require fast response times.
0133One traditional solution to mitigate high peak video streams is to use a technique often referred to as "Constant Bit Rate" (CBR) encoding. The term CBR seems to mean that all frames are compressed to have the same bit rate (ie, size), but that usually refers to a certain number of frames (here, one frame). It is a compression paradigm that allows maximum bitrates over. For example, in the case of Figure 6c, if a CBR constraint is applied to an encoding that limits the bit rate to, for example, 70% of the maximum rated data rate 621, the compression algorithm uses 70% or more of the maximum rated data rate 621. And limit the compression of each frame so that the frames that are normally compressed are compressed with fewer bits. As a result, frames that require more bits than usual to maintain a given quality level are "deficient" in bits, and the video quality of these frames requires more than 70% of the maximum rated data rate of 621. Not worse than for other frames. This solution is acceptable for certain formats of compressed video that (a) have little expected movement or scene change and (b) allow the user to tolerate periodic quality degradation. Can be produced. A good example of a suitable application for CBR is video teleconferencing. This is because if there are only a few peaks and the quality deteriorates for a short time (eg, the camera pans, causing significant scene movement and large peaks, which is sufficient for high quality video compression in the pan. This is because most users will accept it if there are no bits (when there is no bit and it causes poor video quality). Unfortunately, CBR is not well suited for many other applications that have very complex or moving scenes and / or require a reasonably constant level of quality.
0134The short latency compression logic 404 used in one embodiment uses a number of different techniques to address a range of problems associated with streaming short latency compressed video while maintaining high quality. .. First, the short-latency compression logic 404 only generates I-frames and P-frames, alleviating the need to wait for a large number of frames to decode each B-frame. Further, as shown in FIG. 7a, in one embodiment, the short latency compression logic 404 subdivides each uncompressed frame 701-760 into a series of "tiles", and each tile is an I-frame or a P-frame. Encode individually as one of. The compressed I-frame and P-frame groups are referred to herein as "R-frames" 711-770. In the particular example shown in Figure 7a, each uncompressed frame is subdivided into a 16-tile 4x4 matrix. However, these basic principles are not limited to any particular subdivision skim.
0135In one embodiment, the short latency compression logic 404 divides a video frame into a number of tiles and encodes (ie, compresses) one tile from each frame as an I frame (ie, the tiles are full). Compressed as if it were an individual video frame 1/16 of the video size, the compression used for this "mini" frame is an I-frame compression), and the remaining tiles are encoded as P-frames (ie). Compress (ie, the compression used for each 1/16 "mini" frame is a P-frame compression). Tiles that are compressed as I-frames and P-frames are referred to as "I-tiles" and "P-tiles," respectively. For each successive video frame, the tile that should be encoded as the I tile changes. Therefore, at a given frame time, only one of the tiles in the video frame is an I tile, and the remaining tiles are P tiles. For example, in Figure 7a, tile 0 of uncompressed frame 701 is tile I.<sub>0</sub>Encoded as, and the remaining 1-15 tiles are P tiles P<sub>1</sub>-P<sub>15</sub>Encoded as, forming R frame 711. In the next uncompressed video frame 702, tile 1 of uncompressed frame 701 is tile I<sub>1</sub>Encoded as, and the remaining tiles 0 and 2 to 15 are P tiles P<sub>0</sub>And P<sub>2</sub>From P<sub>15</sub>Encoded as, forming an R frame 712. Therefore, the I tile and the P tile as tiles are interleaved one after another in time over consecutive frames. In this process, the last tile in the matrix is the I tile (ie, I)<sub>15</sub>) And continues until R tile 770 is generated. This process is then restarted to generate another R frame, such as frame 711 (ie, encode the I tile for tile 0, and so on). Although not shown in FIG. 7a, in one embodiment, the first R frame of the R frame video sequence contains only I tiles (ie, subsequent P frames are the basis for calculating motion). To have video data). Alternatively, in one embodiment, the startup sequence uses the same I tile pattern as usual, but does not include P tiles for tiles that have not yet been encoded with I tiles. In other words, some tiles are not encoded in the data until the arrival of the first I tile, thus avoiding the startup peak at the video stream data rate 934 in Figure 9a, which is detailed below. In addition, a variety of different sizes and shapes can be used for tiles, while still conforming to these basic principles, as described below.
0136Video decompression logic 412 running on client 415 decompresses each tile as if it were a separate video sequence of small I and P frames, and then renders each tile to framebuffer drive display device 422. For example, I from R frame 711 to 770<sub>0</sub>And P<sub>0</sub>Unzip and render tile 0 of the video footage using. Similarly, I from R frames 711 to 770<sub>1</sub>And P<sub>1</sub>Use to reconstruct tile 1, and so on. As mentioned above, decompressing I-frames and P-frames is a well-known technique, and decompressing I-tiles and P-tiles is accomplished by obtaining multiple instances of video decompression performed on client 415. Can be done. The multiplication process is likely to increase the computational load on the client 415, but in reality it is not. This is because the tile itself is proportionally smaller than the number of additional processes, so the number of pixels displayed is one process and uses traditional full size I and P frames. Because it is the same as the case.
0137This R-frame technique significantly reduces the bandwidth peaks typically associated with the I-frames shown in Figures 6b and 6c. This is because a given frame is usually made up of a P-frame, which is typically smaller than an I-frame. For example, assuming again that a typical I-frame is 160Kb, the I-tile for each frame shown in Figure 7a is approximately 1/16 or 10Kb of this amount. Similarly, assuming a typical P-frame is 16Kb, the P-frame for each tile shown in Figure 7a is approximately 1Kb. The final result is an R frame of approximately 10Kb + 15 * 1Kb = 25Kb. Therefore, each 60-frame sequence is 25Kb * 60 = 1.5Mbps. Thus, at 60 frames per second, this requires a channel that can maintain a bandwidth of 1.5 Mbps, but due to the I tile, significantly lower peaks are distributed over the 60 frame interval.
0138Note that in the previous example assuming the same data rate for I and P frames, the average data rate was 1.1 Mbps. This means that in the previous example, a new I-frame is introduced only once every 60 frames of time, while in this example it creates an I-frame cycle with a time of 16 frames, and thus a time equal to the I-frame16. This is because the tiles are introduced every 16 frames of time, resulting in a significantly higher average data rate. In fact, more frequent I-frames do not increase the data rate linearly. This is due to the fact that the P frame (or P tile) mainly encodes the difference from the previous frame to the next frame. Therefore, if the front frame is quite similar to the next frame, the P frame will be very small, while if the front frame is significantly different from the next frame, the P frame will be very small. growing. However, since P-frames are primarily derived from the previous frame, not the actual frame, the resulting encoded frame contains a sufficient number of bits to contain errors larger than the I-frame (eg, visual defects). be able to. Then, when one P-frame is followed by another, error accumulation occurs, which is exacerbated when there is a long sequence of P-frames. Here, a sophisticated video compressor detects that the quality of the video deteriorates after a series of P-frames, and if necessary, allocates more bits to the subsequent P-frames to improve the quality, or If it is the most efficient course of action, replace the P frame with an I frame. Thus, when a long sequence of P-frames (eg, 59 P-frames as in the example above) is used, typically P-frames, especially when the scene has a great deal of complexity and / or movement. As is further away from the I frame, more bits are needed in the P frame.
0139Alternatively, when looking at the P frame from the opposite perspective, the P frame that follows the I frame tends to require less bits than the P frame that is further away from the I frame. Therefore, in the example shown in FIG. 7a, no P-frame is more than 15 frames away from the preceding I-frame, while in the previous example, the P-frame can be 59 frames away from the I-frame. Therefore, the more often I-frames are, the smaller the P-frames are. Of course, the exact relative size will vary based on the nature of the video stream, but in the example in Figure 7a, if the I tile is 10Kb, the P tile will average only 0.75Kb in size, 10Kb +. At 15 * 0.75Kb = 21.25Kb, or at 60 frames / sec, the data rate is 21.25Kb * 60 = 1.3Mbps, or about 16 from a 1.1Mbps stream with 59 P-frames following an I-frame. % Higher data rate. Again, the relative results between these two solutions for video compression vary based on the video sequence, but typically with R frames than with I / P frame sequences. Experiments have shown that it requires about 20% more bits for a given quality level. But of course, R-frames dramatically reduce peaks and make video sequences available with much shorter latency than I / P frame sequences.
0140R-frames can be configured in a variety of different ways, depending on the nature of the video sequence, the reliability of the channels, and the data rates available. In another embodiment, a different number of tiles than 16 in a 4x4 configuration is used. For example, 2 tiles can be used in a 2x1 or 1x2 configuration, 4 tiles can be used in a 2x2, 4x1 or 1x4 configuration, and 6 tiles can be used in a 3x2, 2x3, 6x1 or 1x6 configuration. It can be used, or 8 tiles can be used in 4x2 (as shown in Figure 7b), 2x4, 8x1 or 1x8 configurations. Note that the tiles do not have to be square, nor do the video frames need to be square or rectangular. The tiles can be divided into shapes that best suit the video frame and application used.
0141In another embodiment, the cycle of I and P tiles is not fixed to the number of tiles. For example, in an 8-tile 4x2 configuration, a 16-cycle sequence can still be used, as shown in Figure 7b. The sequential uncompressed frames 721, 722, and 723 are each divided into 8 tiles, 0-7, and each tile is individually compressed. From R frame 731, only tile 0 is compressed as an I tile and the remaining tiles are compressed as P tiles. For subsequent R frames 732, all eight tiles are compressed as P tiles, then for subsequent R frames 733, tile 1 is compressed as I tiles, and all other tiles are compressed as P tiles. To. Therefore, the sequence continues for 16 frames, I tiles occur only in every other frame, and the last I tile occurs for tile 7 during the 15th frame time (not shown in Figure 7b). ), And during the 16th frame time, the R frame 780 is compressed using all P tiles. The sequence is then restarted with tile 0 being compressed as an I tile and the other tiles being compressed as a P tile. As in the previous embodiment, each first frame of the entire video sequence is typically all I tiles, giving a reference for P tiles forward from that point. The cycle of I tiles and P tiles does not have to be an even multiple of the number of tiles. For example, with eight tiles, each frame with an I tile is followed by two frames with all P tiles, followed by another I tile. In yet another embodiment, some tiles, for example, are known to require more movement and frequent I tiles in one area of the screen, while others are more static (eg,). If you rarely request frequent I tiles (indicating the score of the game), the I tiles are sequenced more often than the other tiles. In addition, each frame is shown in Figure 7a-b with a single I tile, but single (based on the bandwidth of the transmit channel). You can also encode multiple I tiles within a frame of. Conversely, a frame or frame sequence can be transmitted without the I tile (ie, with the P tile only).
0142The reason the solution mentioned in the previous paragraph works well is that not distributing the I tiles across each single frame seems to result in large peaks, but the behavior of the system is not that simple. is there. Each tile is compressed separately from the other tiles, so the smaller the tile, the less efficient the encoding of each tile. This is because the compressor of a given tile cannot take advantage of similar visual features and similar movements from other tiles. Therefore, splitting the screen into 16 tiles is generally less efficient than splitting the screen into 8 tiles. However, if the screen is split into 8 tiles and the data for all I frames is introduced every 8 frames instead of every 16 frames, the overall data rate will be very high. Therefore, introducing all I-frames every 16 frames instead of every 8 frames reduces the overall data rate. Also, by using 8 large tiles instead of 16 small tiles, the overall data rate is reduced and the data peaks caused by the large tiles are reduced to some extent.
0143In another embodiment, the short latency video compression logic 404 of FIGS. 7a and 7b preconfigures the allocation of bits to the various tiles in the R frame with settings based on the known characteristics of the video sequence to be compressed. Control by doing so or automatically based on an ongoing analysis of video quality at each tile. For example, in one competitive video game, the front of the player's car (relatively less moving in the scene) occupies most of the lower half of the scene, while the upper half of the scene is almost always in motion. Completely filled with roads, buildings and landscapes. If the compression logic 404 allocates an equal number of bits to each tile, the tiles in the lower half of the screen of uncompressed frame 721 in Figure 7b (tiles 4-7) are generally the uncompressed frame 721 in Figure 7b. Compressed with higher quality than the tiles in the upper half of the screen (tiles 0-3). If this particular game or this particular scene of the game is known to have such characteristics, the hosting service 210 operator allocates more bits to the tiles at the top of the screen than the tiles at the bottom of the screen. The compression logic 404 can be configured as follows. Alternatively, the compression logic 404 can evaluate the compression quality of the tile after the frame has been compressed (using one or more of a number of compression quality metrics such as peak signal-to-noise ratio (PSNR)). And, over a period of time, if you decide that a tile consistently produces high quality results, then gradually more bits of tiles that produce low quality results until the various tiles reach similar quality levels. Assign to. In another embodiment, the compression logic 404 allocates bits to a particular tile or group of tiles for high quality. For example, it is possible to give an overall good perceptual appearance so that the quality of the center is higher than the edge of the screen.
0144In one embodiment, in order to improve the resolution of an area of the video stream, the video compression logic 404 extracts the area of the video screen where the scene complexity and / or motion is relatively large, the scene complexity and / or motion. Encodes using tiles that are smaller than the area of the relatively small video screen. For example, as shown in Figure 8, small tiles are used around a moving character 805 in one area of one R frame 811 (potentially followed by a series of R frames with the same tile size ( (Not shown) follows). Then, as the character 805 moves into a new area of the image, small tiles are used around this new area in another R-frame 812, as illustrated. As mentioned above, a variety of different sizes and shapes can be used as "tiles" while still conforming to these basic principles.
0145The circular I / P tiles described above substantially reduce the peak data rate of the video stream, but especially for rapidly changing or very complex video footage occurring in video, video games and certain application software. In some cases, the peak cannot be completely eliminated. For example, during a sudden scene transition, a complex frame may be followed by another completely different complex frame. Even if a large number of I tiles may precede the scene transition for only a few frames of time, they do not help in this situation. This is because the new frame material has nothing to do with the I tile in the foreground. In this situation (and other situations where many images change, even if everything doesn't change), the video compressor 404 will code many, if not all, of the P tiles more efficiently as I tiles. As a result, it is determined that a very large peak occurs in the data rate for that frame.
0146As mentioned above, for most consumer grade internet connections (and many office connections), data that exceeds the maximum available data rate shown as 622 in Figure 6c is "jamed" with a rated maximum data rate of 621. There are cases where it is simply not possible to do. The rated maximum data rate of 621 (eg, "6Mbps DSL") is essentially a marketing number for users considering purchasing an Internet connection, but generally does not guarantee a level of performance. For the purposes of the present invention, this is irrelevant. That's because when video is streamed over a connection, only the maximum data rate of 622 available is a problem. As a result, in FIGS. 9a and 9c, when describing a solution to the peak problem, the rated maximum data rate is removed from the graph and only the maximum available data rate 922 is shown. The video stream data rate must not exceed the maximum available data rate of 922.
0147To address this, the first thing the video compressor 404 should do is determine the peak data rate 941, which is the data rate that the channel can steadily handle. This rate can be determined by a number of techniques. One such technique, in Figures 4a and 4b, gradually sends a gradually higher data rate test stream from the hosting service 210 to the client 415, and the client gives feedback to the hosting service regarding packet loss and latency levels. It is something that makes you do. When packet loss and / or latency begins to show a sharp increase, it is an indication that the maximum available data rate of 922 is being reached. The hosting service 210 then gradually reduces the data rate of the test stream until client 415 reports that the test stream is received at a packet loss level that is acceptable for a reasonable period of time and the latency is near minimal. be able to. This establishes a peak maximum data rate of 941, which is then used as the peak data rate for streaming video. Over time, the peak data rate 941 fluctuates (for example, if another user in the house begins to use the Internet connection heavily), and the client 415 constantly monitors it, increasing packet loss or latency. You need to find out if the maximum available data rate 922 is lower than the previously established peak data rate 941, and if so, the peak data rate 941. Similarly, over time, if client 415 finds that packet loss and latency are kept at optimal levels, the video compressor slowly increases the data rate to increase the maximum available data rate. You can request to find out if (for example, if another user in your home ceases heavy use of your internet connection), and the maximum data available. Packet loss and / or long latency indicate that the data rate has exceeded 922, and a lower level can be found again for the peak data rate 941, but it is probably higher than the level before the increased data rate test. Wait again until it's done. Therefore, by using this technique (and other techniques similar to it), the peak data rate 941 can be found and adjusted cyclically as needed. The peak data rate 941 establishes the maximum data rate that can be used by the video compressor to stream video to the user. The logic for determining the peak data rate can be embodied in the user's home 211 and / or hosting service 210. At the user's house 211, the client device 415 performs a calculation to determine the peak data rate and returns that information to the hosting service 210, where at the hosting service 210, the hosting service server 402 receives from the client 415. Calculations are made to determine the peak data rate based on statistical information (eg, packet loss, latency, maximum data rate, etc.).
0148FIG. 9a shows a video stream data rate 934 with substantial scene complexity and / or motion generated using the circular I / P tile compression techniques shown in FIGS. 7a, 7b and 8 described above. Illustrate. Note that the video compressor 404 is configured to output compressed video at an average data rate lower than the peak data rate 941, and most of the time the video stream data rate is kept below the peak data rate 941. A comparison of the video stream data rate 634 and data rate 934 shown in Figure 6c generated using I / P / B or I / P frames produces a very smooth data rate with circular I / P tile compression. Indicates to do. In addition, at frame 2x peak 952 (which approaches 2x peak data rate 942) and frame 4x peak 954 (which approaches 4x peak data rate 944), the data rate exceeds and accepts peak data rate 941. I can't. In fact, even in high-action video from fast-changing video games, peaks with peak data rates above 941 occur in less than 2% of frames, peaks with 2x peak data rates above 942 rarely occur, and 3x peaks. Almost no peaks above the data rate of 943 occur. However, when they occur (eg during scene transitions), the data rates required by them are necessary to produce good quality video footage.
0149One way to solve this problem is to simply configure the video compressor 404 so that the maximum data rate output has a peak data rate of 941. Unfortunately, the quality of the video output it produces during peak frames is inadequate because the compression algorithm is "deficient" in bits. As a result, compression defects appear when there are sudden transitions or fast movements, and eventually the user realizes that defects appear whenever there are sudden changes or rapid movements, which is extremely annoying.
0150The human visual system is extremely sensitive to visual defects that appear during sudden changes or rapid movements, but is less sensitive to the detection of frame rate reductions in such conditions. In fact, when such a sudden change occurs, the human visual system is fascinated by tracking the change, the frame rate drops from 60 fps to 30 fps in a short time, and then immediately 60 fps. You may not notice the return to. And in the case of a very rapid transition, such as a sudden scene switch, the human visual system is unaware that the frame rate drops to 20 fps or 15 fps and then immediately returns to 60 fps. To the observer, the video appears to run continuously at 60 fps, as long as the frame rate drops only occasionally.
0151This property of the human visual system is utilized by the technique shown in Figure 9b. Server 402 (from FIGS. 4a and 4b) produces an uncompressed video output stream at a constant frame rate (60 fps in one embodiment). The timeline shows each frame 961-970 in 1/60 seconds. Each uncompressed video frame starts at frame 961 and is output to a short latency video compressor 404, which compresses the frame in less than one frame time, producing compressed frame 1 981 as the first frame. The data generated for compressed frame 1 981 may be large or small based on a number of factors, as described above. If the data is small enough to be sent to client 415 within the frame time (1/60 seconds) or with a peak data rate of less than 941, it will be sent during the send time (xmit time) 991 (arrows). The length indicates the width of the transmission time). At the next frame time, server 402 generates uncompressed frame 2 962, which is compressed frame 2. Compressed to 982 and transmitted to client 415 during transmission time 992, which is less than frame time at peak data rate 941.
0152Then, at the next frame time, server 402 generates uncompressed frame 3 963. When this is compressed by the video compressor 404, the resulting compressed frame 3 983 is more data than can be transmitted at a peak data rate of 941 in one frame time. Therefore, it is transmitted during the transmit time (2x peak) 993, which takes up the entire frame time and a portion of the next frame time. Now, during the next frame time, the server 402 generates another uncompressed frame 4964 and outputs it to the video compressor 404, but the data is ignored and shown at 974. This is because the video compressor 404 is configured to ignore yet another uncompressed video frame that arrives while still transmitting the previous compressed frame. Of course, the video decompressor of the client 415 does not receive frame 4 and simply keeps displaying frame 3 on the display device 422 for the duration of 2 frames (ie, the frame rate is easily reduced from 60 fps to 30 fps).
0153For the next frame 5, server 402 outputs uncompressed frame 5 965, which is compressed into compressed frame 5 985 and transmitted within one frame of transmission time 995. The video decompressor of client 415 decompresses frame 5 and displays it on display device 422. The server 402 then outputs uncompressed frame 6 966, and the video compressor 404 compresses it into compressed frame 6 986, at which time the data obtained is very long. Compressed frames are transmitted at a peak data rate of 941 during transmission time (4x peak) 996, but it takes approximately 4 frames to transmit the frames. During the next 3 frame time, the video compressor 404 ignores 3 frames from the server 402, and the decompressor of the client 415 steadily holds frame 6 on the display device 422 during the 4 frame time ( That is, the frame rate is easily reduced from 60fps to 15fps). Finally, the server 402 outputs frame 10 970 and the video compressor 404 compresses it frame 10. Compressed to 987, this is transmitted during transmission time 997, and the decompressor of client 415 decompresses frame 10 and displays it on display device 422, and the video resumes at 60 fps again.
0154The video compressor 404 drops the video frame from the video stream generated by the server 402, but does not drop the audio data regardless of how the audio arrives, but the audio data when the video frame is dropped. Note that it continues to compress and sends them to the client 415, which continues to decompress the audio data and feeds the audio to the device used by the user to play the audio. Therefore, the audio remains unattenuated for the duration of the frame drop. Compressed audio consumes a relatively small percentage of bandwidth compared to compressed video, and as a result does not significantly affect the overall data rate. Although not shown in any of the data rate diagrams, the data rate capacity is always reserved for the compressed audio stream within the peak data rate 941.
0155The example described above for Figure 9b was chosen to show how the frame rate drops during data rate peaks, but so when the circular I / P tile technique described above is used. It has not been shown that high data rate peaks and the frames dropped thereby are rare even in high scene complexity / high action sequences such as those that occur in video games, videos and certain application software. Therefore, the reduction in frame rate is occasional and short-lived, which the human visual system does not detect.
0156When the frame rate reduction mechanism described above is applied to the video stream data rate shown in FIG. 9a, the resulting video stream data rate is shown in FIG. 9c. In this example, the 2x peak 952 is reduced to a flat 2x peak 953, the 4x peak 955 is reduced to a flat 4x peak 955, and the total video stream data rate 934 is kept below the peak data rate 941.
0157Thus, the techniques described above can be used to transmit high action video streams over common internet and consumer grade internet connections with short latency. In addition, in an office environment on a LAN (eg 100Mbs Ethernet or 802.11g wireless) or a private network (eg 100Mbps connection between a data center and an office), high action video streams can be transmitted without peaks. Allows a large number of users (eg, transmitting 1920x1080 at 60fps at 4.5Mbps) to use a LAN or shared private data connection without overwhelming overlapping peaks on the network or network switch backplane.
0158<u style="single">Data rate adjustment</u> In one embodiment, the hosting service 210 first accesses the maximum available data rate 622 and the latency of the channel to determine an appropriate data rate for the video stream, and then responds to the data rate. Dynamically adjust. To adjust the data rate, the hosting service 210 can change, for example, the resolution of the video and / or the number of frames / sec of the video stream sent to the client 415. The hosting service can also adjust the quality level of the compressed video. When changing the resolution of the video stream, for example from 1280x720 resolution to 640x360, the video decompression logic 412 of the client 415 can scale the video to maintain the same video size on the display screen.
0159In one embodiment, the hosting service 210 pauses the game in a situation where the channel drops out completely. In the case of a multiplayer game, the hosting service reports to the other user that the user has dropped out of the game and / or suspends the game to the other user.
0160<u style="single">Dropped or delayed packets</u> In one embodiment, due to packet loss between the video compressor 404 and the client 415 in FIGS. 4a or 4b, or because the packets are received out of order, it is too late to decompress and the decompression frame latency requirement. Video decompression logic 412 can mitigate visual flaws in case of data loss due to arrival too late to satisfy. In streaming I / P frame realization, if there are lost / delayed packets, it affects the entire screen and potentially freezes the screen completely for a period of time, or other screen width visuals. Display the above defect. For example, if a lost / delayed packet causes a loss of an I-frame, the decompressor will lack a reference for all subsequent P-frames until a new I-frame is received. If the P-frame is lost, it will affect the P-frame for all subsequent screens. Based on how long it takes for an I-frame to appear, this can be a long or short visual effect. With the interleaved I / P tiles shown in Figures 7a and 7b, lost / delayed packets are very unlikely to affect the entire screen. This is because it only affects the tiles contained in the affected packet. If the data for each tile is sent within an individual packet, it will only affect one tile if the packet is lost. Of course, the width of the visual defect depends on whether the I tile packet is lost and, if the P tile is lost, how many frames it will take for the I tile to appear. However, if different tiles on the screen are updated frequently (potentially frame by frame) in I-frames, then even if one tile on the screen is affected, the other tiles are not. In addition, if an event causes the loss of many packets at once (eg, a spike in power following a DSL line can cause a short data flow). Some tiles are affected more than others), but some are only affected for a short time as they are updated quickly with new I tiles. Also, in streaming I / P frame realization, the I frame is not only the most important frame, but also a very large frame, so if there is an event that causes a dropped / delayed packet, The I-frame is more likely to be affected than a very small I-tile (ie, if any I-frame is lost, there is no chance that the I-frame can be decompressed). For all these reasons, using I / P tiles causes far less visual flaws when packets are dropped / delayed than in I / P frames.
0161In one embodiment, attempts are made to reduce the effect of lost packets by intelligently packaging compressed tiles within TCP (Transmission Control Protocol) or UDP (User Datagram Protocol) packets. For example, in one embodiment, tiles are aligned at packet boundaries, if possible. Figure 10a shows how tiles are packed into a series of packets 1001-1005 without embodying this feature. More specifically, in FIG. 10a, tiles intersect packet boundaries and are packed inefficiently so that loss of a single packet results in loss of multiple frames. For example, if packets 1003 or 1004 are lost, three tiles will be lost, leading to visual flaws.
0162In contrast, FIG. 10b shows the tile packing logic 1010 for intelligently packing tiles in a packet to reduce the effects of packet loss. First, the tile packing logic 1010 aligns the tiles at the packet boundaries. Therefore, tiles T1, T3, T4, T7 and T2 are each aligned to the boundary of packet 1001-1005. The tile packing logic also attempts to fit tiles within a packet in the most efficient way possible, without crossing the packet boundaries. Based on the size of each tile, tiles T1 and T6 are combined in one packet 1001, T3 and T5 are combined in one packet 1002, and tiles T4 and T8 are combined in one packet 1003. Tile T8 is added to packet 1004, and tile T2 is added to packet 1005. Therefore, under this skim, the loss of a single packet results in the loss of two or less tiles (rather than the three tiles shown in Figure 10a).
0163One additional benefit to the embodiment shown in Figure 10b is that the tiles are transmitted in a different order than they appear in the video. Thus, when adjacent packets are lost from the same event that interferes with transmission, they affect non-neighboring areas on the screen and make defects on the display less noticeable.
0164In one embodiment, forward error correction (FEC) techniques are used to protect some parts of the video stream from channel errors. As is known in this technique, FEC techniques such as Reed-Solomon and Viterbi generate error correction data information and attach it to the data transmitted over the communication channel. If an error occurs in the underlying data (eg, I-frame), FEC can be used to correct the error.
0165FEC codes increase the data rate of transmission, so ideally they are only used where they are most needed. If data is transmitted that does not cause very noticeable visual defects, it is preferable not to use FEC code to protect the data. For example, the P tile immediately before the lost I tile causes a visual defect on the screen for 1/60 second (ie, the tile on the screen is not updated). Such visual defects can be barely detected by the human eye. If the P tile is further back from the I tile, losing the P tile will make it even more noticeable. For example, if the tile cycle pattern is that the I tile is followed by 15 P tiles, and then the I tile is obtained again, if the P tile immediately following the I tile is lost, that tile will be 15 It will show the wrong image for the duration of the frame (at 60fps, that's 250ms). The human eye can easily detect stream interruptions for 250 ms. Therefore, the further back the P tile is from the new I tile (ie, the closer the P tile follows the I tile), the more noticeable the defects will be. As mentioned above, in general, the closer the P tile follows the I tile, the smaller the data for that P tile. Therefore, not only is it more important to protect the P tiles that follow the I tiles from being lost, but it is also important that they are small in size. And, in general, the smaller the data that needs to be protected, the smaller the FEC code needed to protect it.
0166Therefore, as shown in FIG. 11a, in one embodiment, due to the importance of the I tile in the video stream, only the I tile is given the FEC code. Therefore, FEC1101 contains an error correction code for I tile 1100, and FEC1104 contains an error correction code for I tile 1103. In this embodiment, no FEC is generated for the P tile.
0167In one embodiment shown in Figure 11b, the FEC code is also generated for P tiles that, if lost, are most likely to cause visual defects. In this embodiment, FEC1105 gives error correction codes for the first three P tiles, but not for subsequent P tiles. In another embodiment, the FEC code is generated for the P tile with the smallest data size (it tends to occur the earliest after the I tile and self-select the P tile that is most important to protect).
0168In another embodiment, instead of transmitting the FEC code with the tile, the tile is transmitted twice, each time in a different packet. If one packet is lost / delayed, the other packet is used.
0169In one embodiment shown in FIG. 11c, FEC codes 1111 and 1113 are generated for audio packets 1110 and 1112 transmitted from the hosting service at the same time as the video, respectively. Maintaining audio integrity in the video stream is especially important. This is because distorted audio (eg, ticking or soothing) leads to a particularly unwanted user experience. The FEC code helps ensure that the audio content is rendered distortion-free on the client computer 415.
0170In another embodiment, instead of transmitting the FEC code with the audio data, the audio data is transmitted twice, in different packets each time. If one packet is lost / delayed, the other packet is used.
0171Further, in one embodiment shown in FIG. 11d, FEC codes 1121 and 1123 are used for user input commands 1120 and 1122 (eg, button presses) transmitted upstream from the client 415 to the hosting service 210, respectively. This is important. This is because the lack of button press or mouse movement in a video game or application leads to an unwanted user experience.
0172In another embodiment, instead of transmitting the FEC code with the user input command data, the user input command data is transmitted twice, each time in a different packet. If one packet is lost / delayed, the other packet is used.
0173In one embodiment, the hosting service 210 evaluates the quality of the communication channel with the client 415 to determine if FEC should be used, and if so, which part of the video, audio and user commands FEC is applied. Decide if it should be. Evaluating the "quality" of a channel includes functions such as evaluating packet loss, latency, etc., as described above. If the channel is not particularly reliable, the hosting service 210 can apply FEC to all I tiles, P tiles, audio and user commands. In contrast, if the channel is reliable, the hosting service 210 applies FEC only to audio and user commands, or FEC to audio or video, or does not use FEC at all. Various other permutations of FEC application can be used, still in line with these basic principles. In one embodiment, the hosting service 210 constantly monitors the status of the channel and changes the FEC policy accordingly.
0174In another embodiment, referring to Figures 4a and 4b, the FEC will lose tile data due to packet loss / delay and loss of tile data, or perhaps due to particularly bad packet loss. If it cannot be fixed, the client 415 evaluates how many frames are left before the new I tile is received and compares it to the round-trip latency from the client 415 to the hosting service 210. If the round-trip wait time is less than the number of frames before a new I-tile arrives, the client 415 sends a message requesting the new I-tile to the hosting service 210. This message is routed to a video compressor 404, which generates an I tile instead of a P tile for the tile where the data was lost. If the system shown in Figures 4a and 4b is typically designed to give a round trip latency of less than 80ms, it will fix the tile within 80ms (at 60fps, the frame will be 16.67ms). Therefore, in the time of all frames, the waiting time of 80ms is the time of 5 frames 83. The tiles are modified within 33ms, i.e., this is a noticeable break, but much less noticeable than, for example, a 250ms break for 15 frames). When the compressor 404 generates I tiles from its normal circulation order, if the I tiles exceed the available bandwidth of that frame, the compressor 404 delays the cycles of the other tiles, Allow other tiles to receive P tiles during that frame time (even if one tile should normally be an I tile in that frame), then start at the next frame and continue the normal cycle , And make sure that the tile that normally receives the I tile in the frame before it receives the I tile. This action delays the phase of the R-frame circulation for a short time, but is usually not visually noticeable.
0175<u style="single">Video and audio compressor / decompressor realization</u> FIG. 12 shows one particular embodiment of compressing eight tiles in parallel using a multicore and / or multiprocessor 1200. In one embodiment, a dual processor quad-core Xeon CPU computer system operating above 2.66GHz is used, with each core embodying an open source x264 H.264 compressor as an independent process. However, various other hardware / software configurations may be used while still conforming to these underlying principles. For example, each of the CPU cores can be replaced with an H.264 compressor embodied in FPGA. In the example shown in Figure 12, core 1201-1208 is used to process the I and P tiles simultaneously as eight independent threads. As is well known in this technology, today's multi-core and multi-processor computer systems are essentially Microsoft Windows. Multithreading is possible when integrated with multithreaded operating systems such as XP Professional Edition (64-bit or 32-bit edition) and Linux.
0176In the embodiment shown in FIG. 12, each of the eight cores serves only for one tile, so it operates primarily independently of the other cores that each perform individual instantiation of x264. Capture uncompressed video at 640x480, 800x600, or 1280x720 resolution using a PCI Express x1-based DVI capture card, such as the Sendero Video Imaging IP Development Board from Microtronics in Austin, Netherlands. And the FPGA on this card uses direct memory access (DMA) to transfer the captured video over the DVI bus to system RAM. The tiles are arranged in a 4x2 array 1205 (they are shown as square tiles, but in this embodiment they have a resolution of 160x240). Each x264 instantiation is configured to compress one of eight 168x240 tiles, which, after the first I tile compression, each core goes into a cycle and each one frame is the other frame. And are synchronized as shown in Figure 12 to compress one I tile and then seven P tiles.
0177At each frame time, the resulting compressed tiles are combined into a packet stream using the techniques described above, and the compressed tiles are then sent to the destination client 415.
0178Although not shown in Figure 12, if the data rate of the eight tiles to be combined exceeds a certain peak data rate of 941, it is necessary until the data of the eight tiles to be combined is transmitted. A total of 8x264 processes are suspended for a number of frames, depending on the number of frames.
0179In one embodiment, the client 415 is embodied as software on a PC running eight instances of FFmpeg. The receiving process receives eight tiles, and each tile is routed to FFmpeg, which decompresses the tiles and renders them at the appropriate tile positions on the display device 422.
0180Client 415 receives keyboard, mouse, or game controller input from the PC's input device driver and sends it to server 402. The server 402 then applies the received input device data to a game or application running on the server 402, which is a PC running Windows using an Intel 2.16GHz Core Duo CPU. Server 402 then generates a new frame and outputs it through its DVI output, from a motherboard-based graphics system, or through the DVI output of an NVIDIA 8800 GTX PCI Express card.
0181At the same time, server 402 outputs the audio formed by the game or application through its digital audio output (eg S / PDIF), which is the digital audio input of a dual quad-core Xeon-based PC that embodies video compression. Combined with. The Vorbis open source audio compressor is used to compress audio at the same time as video, using the cores available for process threads. In one embodiment, the core that completes the tile compression performs the audio compression first. The compressed audio is then transmitted with the compressed video and decompressed on the client 415 using the Vorbis audio decompressor.
0182<u style="single">Hosting service server center distribution</u> Light through a glass, such as an optical fiber, travels at a portion of the speed of light in a vacuum, thus allowing the exact propagation speed of light in an optical fiber. However, in practice, given the time of routing delays, transmission inefficiencies, and other overheads, it is observed that the optimal latency on the Internet reflects a transmission rate of nearly 50% of the speed of light. .. Therefore, the optimal 1000 mile round trip wait time is about 22 ms, and the optimal 3000 mile round trip wait time is about 64 ms. Therefore, a single server on one coast of the United States is too far away to serve clients on the other coast (about 3000 miles away) with the desired latency. However, as shown in Figure 13a, the server center 1300 for hosting service 210 is located in the central United States (eg, Kansas, Nebraska, etc.) and is less than or equal to about 1500 miles to any point in the Americas. In that case, the round-trip Internet waiting time can be reduced to about 32 ms. With reference to Figure 4b, the worst-case latency allowed for user ISP453 is typically 25ms, but DSL and cable modem systems have observed latency close to 10-15ms. Please note. Figure 4b also assumes that the maximum distance from the user's house 211 to the hosting center 210 is 1000 miles. Therefore, a typical user ISP round-trip latency of 15 ms is used, and at a maximum internet distance of 1500 miles for a round-trip latency of 32 ms, the point at which the user operates input device 421 and sees a response on display device 422. The total round-trip waiting time from is 1 + 1 + 15 + 32 + 1 + 16 + 6 + 8 = 80ms. Therefore, a response time of 80 ms is typically obtained over an internet distance of 1500 miles. This allows a user's home with a short enough user ISP wait time 453 in the Americas to access a single centrally located server.
0183In another embodiment shown in Figure 13b, the server centers HS-1 to HS6 of the hosting service 210 are strategically located in the United States (or other geographic area) and have one large hosting service server center (eg HS2). And HS5) will be located near the populous center. In one embodiment, the server centers HS1 to HS6 exchange information via the Internet and / or a combination of private networks, Network 1301. With multiple server centers, it is possible to provide services to users with a long user ISP wait time 453 with a short wait time.
0184Distance on the Internet is certainly a factor that contributes to round-trip latency through the Internet, but from time to time other factors that are less relevant to latency come into play. From time to time, packet streams are routed distantly through the Internet and back again, leading to latency from long loops. Occasionally, there are routing devices on the path that do not work properly, leading to transmission delays. Occasionally there is overloaded traffic on the route, introducing delays. And from time to time, failures occur that completely prevent the user's ISP from being routed to a given destination. Therefore, the general Internet usually has a very reliable optimal route and latency, which is largely determined by distance (especially for long-distance connections that give rise to routes outside the user's local area). It gives a connection from one point to another, but such reliability and latency are not guaranteed and often get from the user's home to a given destination on the general Internet. Can't.
0185In one embodiment, when the user client 415 first connects to the hosting service 210 to play a video game or use an application, the client is available at the hosting service server center HS1 to HS6 at startup. Communicate with each (eg, using the techniques described above). If the latency is short enough for a particular connection, that connection is used. In one embodiment, the client communicates with all or a subset of the hosting service server centers where the one with the shortest latency connection is selected. The client may select a service center with the shortest latency connection, or the service center may identify the one with the shortest time connection and give this information to the client (eg, in the form of an internet address). Good.
0186If a particular hosting service server center is overloaded and / or the user's game or application can tolerate waiting time for another less loaded hosting service server center, the client 415 may use the other hosting service. Directed to the server center. In such a situation, the game or application running by the user is dormant on server 402 in the server center due to the user's overload condition, and the state data of the game or application is stored on the server in another hosting service server center. Transferred to 402. The game or application is then restarted. In one embodiment, the hosting service 210 waits until the game or application reaches a natural pause (eg, between levels in the game, or after the user initiates a "save" operation in the application) and makes a transfer. To do. In yet another embodiment, the hosting service 210 waits for the user's activity to pause for a period of time (eg, 1 minute), at which point the transfer begins.
0187As mentioned above, in one embodiment, the hosting service 210 contracts with the Internet bypass service 440 of FIG. 14 in an attempt to give its clients a guaranteed latency. The Internet bypass service used here is a service that provides a private network route from one point to another on the Internet with guaranteed characteristics (eg, latency, data rate, etc.). For example, if hosting service 210 receives a large amount of traffic from users using AT & T's DSL service provided in San Francisco instead of routing to AT & T's San Francisco-based central office, hosting service 210 will be in San Francisco. You can lease large amounts of private data connections from a service provider (perhaps AT & T itself or another provider) between your base central office and one or more service centers for hosting service 210. Then, if the route from all hosting service server centers HS1 to HS6 to San Francisco users using AT & T's DSL over the general Internet results in very long latency, then use a private data connection instead. Can be used. Private data connections are generally more expensive than routes through the general Internet, but as long as they are kept at a small percentage of the hosting service 210 connections to the user, the impact on the total cost is low and the user Will experience a more consistent service.
0188Service centers often have two layers of backup power in the event of a power outage. The first layer is typically a backup power source from the battery (or from another readily available energy source, such as a flywheel that is attached to the generator and kept running), with the main power source. Power immediately in the event of a power outage to keep the server center running. If the power outage is short and the mains recover quickly (eg, within a minute), then all batteries are needed to keep the server center running. However, if the power outage is long-term, the generator (eg, Ziesel urging) can typically be started to take over the battery and operate as long as there is fuel. Such generators are very expensive because they must be able to generate as much power as the server center would normally get from the mains.
0189In one embodiment, each of the hosting services HS1 through HS5 shares user data with each other, suspends ongoing games and applications in the event of a server center outage, and then provides state data for the games or applications. Transfer from each server 402 to server 402 in another server center, and notify each user's client 415 to instruct them to communicate with the new server 402. If this situation does not occur frequently, it is possible to accept the transfer of the user to the hosting service server center, which cannot provide the optimum latency (ie, the user simply waits a long time during a power outage). (Just allow time), this allows a very wide range of options for transferring users. For example, given a time zone difference across the United States, East Coast users go to bed at 11:30 PM, while West Coast users at 8:30 PM are beginning to peak their use of video games. In the event of a power outage at the West Coast Hosting Service Server Center at that time, other Hosting Service Server Centers may not have enough West Coast Server 402s to handle all users. In such a situation, users at the East Coast hosting service server center with available servers 402 can be transferred, resulting in long wait times only for those users. When a user is transferred from a server center that has lost power, the server center then initiates an orderly shutdown of its servers and equipment, shutting down all equipment before the battery (or other immediate power backup) runs out. can do. In this way, the cost of the generator for the server center can be avoided.
0190In one embodiment, during times of heavy load on hosting services 210 (either due to peak user load or due to one or more server centers power outages), users wait for the game or application they are using. Transferred to another server center based on time requirements. Thus, users using games or applications that require low latency are given priority over available short latency server connections when power is limited.
0191<u style="single">Features of hosting service</u> FIG. 15 shows an embodiment of a server center component for the hosting service 210 used in the following feature description. Similar to the hosting service 210 shown in FIG. 2a, the components of the server center are controlled and aligned by the control system 401 of the hosting service 210 unless otherwise specified.
0192Inbound Internet traffic 1501 from user client 415 is directed to inbound routing 1502. Typically, inbound Internet traffic 1501 enters the server center via a high-speed fiber optic connection to the Internet, but adequate bandwidth, reliability, and short latency network connectivity are sufficient. The inbound routing 1502 is a network (the network can be embodied as an Ethernet network, a fiber channel network, or via other means of transport) switches, and a system of routing servers that support the switches, which allow incoming packets. Take up and route each packet to the appropriate application / game server 1521-1525. In one embodiment, packets delivered to a particular app / game server represent a subset of the data received from the client and / or other components within the data center (eg, network components such as gateways and routers). ) May be converted / changed. In some cases, for example, when a game or application runs once in parallel on multiple servers, packets are routed to two or more servers 1521-1525 at a time. The RAID array 1511-1512 is connected to the inbound routing network 1502, allowing the app / game server 1521-1525 to read and write to those RAID arrays 1511-1512. In addition, a RAID array 1515 (which can be embodied as multiple RAID arrays) is also connected to the inbound routing 1502, and data from the RAID array 1515 can be read from the app / game server 1521-1525. Inbound Routing 1502 is a tree structure of switches rooted in inbound Internet traffic 1501, all A wide range of traditional network architectures, including mesh structures that interconnect various devices, or a set of subnets that are interconnected so that concentrated traffic between interconnect devices is separated from concentrated traffic between other devices. It can be embodied in char. A form of network configuration is SAN, which is typically used as a storage device, but can also be used for common high-speed data transfer between devices. Also, each app / game server 1521-1525 has multiple network connections to inbound routing 1502. For example, the server 1521-1525 can have a network connection to a subnet attached to a RAID array 1511-1512 and another network connection to a subnet attached to another device.
0193The app / game server 1521-1525 may be configured to be all the same, some different, or all different, as described above for the server 402 in the embodiment shown in FIG. 4a. In one embodiment, each user typically uses at least one app / game server 1521-1525 when using the hosting service. For the sake of brevity, we will assume that a given user uses the app / game server 1521, but one user can use multiple servers and multiple users can use a single app / game. Server 1521-1525 can be shared. As mentioned above, the user control input sent by the client 415 is received as inbound internet traffic 1501 and routed to the app / game server 1521 through inbound routing 1502. The app / game server 1521 uses the user's control input as control input to a game or application running on the server, and calculates the next frame of video and associated audio. The app / game server 1521 then outputs the uncompressed video / audio 1529 to the shared video compression 1530. The app / game server can output uncompressed video via a means that includes a 1 Gigabit or higher Ethernet connection, but in one embodiment the video is output via a DVI connection and audio and other compressed and communication channels. The status information is output via a universal serial bus (USB) connection.
0194Shared video compression 1530 compresses uncompressed video and audio from the app / game servers 1521-1525. Compression may be embodied entirely in hardware or in hardware execution software. There may be a dedicated compressor for each app / game server 1521-1525, or if the compressor is fast enough, using a given compressor, from two or more app / game servers 1521-1525 Video / audio can be compressed. For example, at 60fps, the video frame time is 16.67ms. If the compressor can compress the frame in 1ms, it can be used to compress video / audio from about 16 app / game servers 1521-1525, which is the input from one server. Is taken after another server, the compressor saves the state of each video / audio compression process, and switches the context as it circulates between the video / audio streams from the server. The result is substantial cost savings in compressed hardware. In one embodiment, the compressor resource is in shared pool 1530 with shared storage means (eg RAM, flash) to store the state of each compression process, because different servers complete frames at different times, and server 1521. When the -1525 frame is complete and ready for compression, the control means decides which compression resource is available at that time and tells the compression resource the state of the server's compression process and the uncompressed video to be compressed. Gives a frame of audio.
0195A portion of the state of the compression process on each server is information about the compression itself, such as the decompression frame buffer data of previous frames that can be used as a reference for P tiles, the resolution of the video output; the quality of compression; the tile structure; per tile. Note that bit allocation; includes compression quality, audio formats (eg, stereo, surround sound, Dolby®, AC-3). However, the state of the compression process is the communication channel state information for the peak data rate 941, whether the previous frame (shown in Figure 9b) is currently output (as a result, whether the current frame should be ignored), and the potential. Also includes channel characteristics that should be considered for compression, such as whether there is excessive packet loss (eg, with respect to I tile frequency, etc.) that affects compression decisions. App / game server 1521 when peak data rate 941 or other channel characteristics change over time, as determined by the app / game server 1521-1525 that supports each user monitoring the data sent by client 415. The -1525 sends relevant information to the shared hardware compression 1530.
0196Also, the shared hardware compression 1530 uses the means described above to packetize the compressed video / audio, and if appropriate, apply FEC code, copy some data, or otherwise. To ensure sufficient ability to receive the video / audio data stream by the client 415 and decompress it with the highest quality and reliability possible.
0197Some applications, such as those described below, require that the video / audio output of a given app / game server 1521-1525 be obtained simultaneously in multiple resolutions (or in multiple other formats). Therefore, if the app / game server 1521-1525 notifies the resources of the shared hardware compression 1530, the uncompressed video / audio 1529 of the app / game server 1521-1525 will have different formats, different resources, and / or Compressed simultaneously with different packet / error correction structures. In some cases, several compression resources can be shared between multiple compression processes that compress the same video / audio (for example, many compression algorithms have multiple sizes of video before applying compression. There is a step to scale to. If you need to output different sized footage, you can use this step to accommodate multiple compression processes at once). In other cases, separate compression resources are required for each format. In any case, all the various resolutions and formats of compressed video / audio 1539 required for a given app / game server 1521-1525 will be outbound routing (one or many). It is output to 1540 at once. In one embodiment, the output of the compressed video / audio 1539 is in UDP format and is therefore a positional stream of packets.
0198The outbound routing network 1540 passes each compressed video / audio stream through the outbound Internet traffic 1599 interface, which is typically connected to a fiber interface to the Internet, by the intended user (s). Or with a set of routing servers and switches directed to other destinations and / or back to the delay buffer 1515 and / or back to inbound routing 1502 and / or through a private network (not shown) for video distribution. ing. Note that the outbound routing 1540 (as described below) can output a given video / audio stream to multiple destinations at once. In one embodiment, it broadcasts a given UDP stream intended to be streamed to multiple destinations at once, and this broadcast is repeated by a routing server and switch in outbound routing 1540. Embodied using (IP) multicast. Multiple destinations for the broadcast are clients 415 for multiple users over the Internet, multiple app / game servers 1521-1525 via inbound routing 1502, and / or one or more delay buffers 1515. Thus, the output of a given server 1521-1522 is compressed into one or more formats, and each compressed stream is directed to one or more destinations.
0199In yet another embodiment, multiple app / game servers 1521-1525 are used simultaneously by a single user (eg, in a parallel processing configuration to produce 3D output for complex scenes) and each server is completed. If you generate a portion of the video, you can combine the video output of multiple servers 1521-1525 into composite frames with shared hardware compression 1530, and from that point on, it's a single app / game. Treated as described above, as if it came from server 1521-1525.
0200In one embodiment, the copy of all video generated by the app / game server 1521-1525 (at least the resolution of the video viewed by the user or higher) is at least a few minutes (15 minutes in one embodiment), delay buffer 1515. Note that it is recorded in. This allows each user to "rewind" the video from each session to review their previous work or performance (in the case of a game). Thus, in one embodiment, each compressed video / audio output 1539 stream routed to the user client 415 is also multicast to the delay buffer 1515. When the video / audio is stored in the delay buffer 1515, the directory in the delay buffer 1515 finds the network address of the app / game server 1521-1525, which is the source of the delayed video / audio, and the delayed video / audio. Gives a cross-reference to and from a position on the delay buffer 1515 that can be.
0201<u style="single">A live, instantly visible, instantly playable game</u> The app / game server 1521-1525 can be used to run a given application or video game for the user, as well as a user interface for the hosting service 210 that supports navigation by the hosting service 210 and other features. It can also be used to generate applications. A screenshot of one such user interface application is shown in Figure 16, the "Game Finder" screen. This particular user interface screen allows the user to view 15 games that are being played live (or delayed) by other users. Each "thumbnail" video window, such as the 1600, is a live, moving video window showing one video from one user's game. The field of view shown in the thumbnail may be the same field of view that the user is looking at, or it may be a delayed field of view (for example, when the user plays a combat game, the user can see where he or she is hiding from other users. You may choose to delay your gameplay horizons for a period of time, eg, 10 minutes). Further, the field of view may be the camera field of view of the game, which is different from the field of view of the user. Through menu choices (not shown in this figure), the user can choose which games to watch at one time based on various criteria. As a small example of exemplary selection, the user can randomly select a game (as shown in Figure 16), all games of one type (all played by different players), only the top ranked players of the game, of the game. You can choose a player at a given level, a player with a lower rank (eg, if the player is learning the basics), a player who is a "buddy" (or rival), a game with the most viewers, and so on.
0202In general, each user determines if the video from their game or application will be seen by others, and if so, what others, when others will see it, and whether it will only be seen late. Note the judgment.
0203The app / game server 1521-1525, which generates the user interface screen shown in Figure 16, gets 15 video / audio feeds by sending a message to the app / game server 1521-1525 for each user requesting a game. To do. This message is sent through inbound routing 1502 or another network. This message includes the requested video / audio size and format and identifies the user viewing the user interface screen. A given user chooses a "privacy" mode and chooses not to allow other users to view the video / audio of his game (from his or another perspective), or in the paragraph above. As mentioned, the user allows to watch the video / audio from his game, but may choose to delay the video / audio to watch. The user app / game server 1521-1525, which receives and accepts the request to allow viewing of video / audio, sends a confirmation to the requesting server and attaches it to the shared hardware compression 1530 in the requested format or screen size. It also notifies the request to generate a compressed video stream (assuming the format and screen size are different from those already generated), and the destination of the compressed video (ie, the requesting server). Also instruct. If the requested video / audio is only delayed, the requesting app / game server 1521-1525 is notified so and is delayed with the location of the video / audio in the directory of the delay buffer 1515. Get the delayed video / audio from the delay buffer 1515 by looking up the network address of the app / game server 1521-1525, which is the source of the video / audio. Once all of these requests have been generated and processed, a raw thumbnail-sized video stream of up to 15 will be outbowed. It is routed from the routing 1540 to the inbound routing 1502 to the app / game server 1521-1525, which raises the user interface screen and is unzipped and displayed by the server. The delayed video / audio stream may have too large a screen size, and if so, the app / game server 1521-1525 decompresses the stream and scales the video stream to thumbnail size. In one embodiment, a request for audio / video is sent to (and managed by) a central "management" service (not shown in Figure 15) similar to the hosting service control system of Figure 4a. Then redirect the request to the appropriate app / game server 1521-1525. Further, in one embodiment, the thumbnail is "pushed" to the client of the user who allows it, so that the request may not be issued.
0204Audio from 15 games, all mixed at the same time, can make an unpleasant sound. The user may choose to mix all the sounds together in this way (perhaps just to get a sense of the noise caused by all the actions they are watching), or the user may choose one game at a time. You may choose to listen only to the audio from. Single game selection is achieved by moving the yellow selection box 1601 to a given game (moving the yellow box can be done using the arrow keys on the keyboard, moving the mouse, or the joystick. Can be achieved by moving the mouse or pressing the direction button on another device such as a mobile phone). If a single game is selected, only the audio from that game will be played. Also, game information 1602 is shown. In the case of this game, for example, the publisher logo (EA) and the game logo Need for Speed The "Carbon", and the orange horizontal bar, indicate in relative terms the number of people playing or watching the game at that particular moment (in this case, many, so the game is "hot". ). In addition, there is also "Stats", which has 145 players Need for Speed. It is actively playing 80 different instantiations of the Game (ie, it can be played by either a personal player game or a multiplayer game), and has 680 viewers (this user is one of them). It shows that it is a person). These statistics (and other statistics) keep a log of the operations of the hosting service 210, as well as bill the users as appropriate and pay the publishers who provide the content, the hosting service control system 401. Collected by and stored in RAID array 1511-1512. Some of the statistics are recorded by the action by the service control system 401, and some are reported to the service control system 401 by the individual app / game server 1521-1525. For example, the app / game server 1521-1525 running this "game finder" application sends a message to the hosting service control system 401 when watching (and finishing) the game, and how many games Allows you to update statistics on whether is being viewed. Some of the statistics are used for user interface applications such as this "Game Finder" application.
0205When the user clicks the activation button on the input device, they can zoom in on the thumbnail video in the yellow box while keeping it raw to full screen size. This effect is shown in the process of Figure 17. Note that the video window 1700 is increasing in size. To embody this effect, the app / game server 1521-1525 is a full-screen video stream of the game routed from the app / game server 1521-1525 running the selected game. Request a copy (at the resolution of the user's display device 422). The app / game server 1521-1525 running the game no longer needs a copy of the game's thumbnail size on the shared hardware compressor 1530 (another app / game server 1521-1525 does not require such a thumbnail. Notify that (as long as) and then instruct to send a full-screen size copy of the video to the app / game server 1521-1525 that is zooming the video. The user playing the game may or may not have a display device 422 with the same resolution as the user zooming in on the game. In addition, others watching the game may or may not have a display device 422 with the same resolution as the user zooming in on the game (and different audio playback means, such as stereo or surround sound. May have). Therefore, the shared hardware compressor 1530 determines if a properly compressed video / audio stream has already been generated that meets the requirements of the user requesting the video / audio stream, and if it exists. Notifies outbound routing 1540 to route a copy of the stream to the app / game server 1521-1525 that is zooming the video, and if not Instruct outbound routing to compress another copy of the video suitable for that user and send the stream back to the app / game server 1521-1525 and inbound routing 1502 that are zooming the video. The server, which is currently receiving the full-screen version of the selected video, unzips it and gradually scales it up to full size.
0206FIG. 18 shows what the screen looks like after the game is fully zoomed in to full screen, where the game is at full resolution of the user's display device 422, as indicated by the image pointed to by arrow 1800. It is indicated by. The app / game server 1521-1525 running the GameFinder application sends a message to another app / game server 1521-1525 that has generated a thumbnail that is no longer needed, as well as hosting where no other game is seen anymore. Send a message to the service control server 401. In this regard, the display generated is an overlay 1801 at the top of the screen, which gives the user information and menu control. Note that as the game progressed, the audience grew to 2503 viewers. Thus, for a large number of viewers, a large number of viewers are tied to a display device 422 with the same or nearly the same resolution (each app / game server 1521-1525 has the ability to scale the video to adjust suitability. Has).
0207Since the illustrated game is a multiplayer game, the user can decide to join the game at some point. Hosting service 210 may or may not allow users to join the game for a variety of reasons. For example, the user may have to pay to play the game, or may choose not to play, and the user may not have sufficient ranking to join the particular game ( For example, it may not be able to withstand competition from other players, or the user's internet connection may not have a short enough wait time for the user to play (eg, wait time constraints for watching the game). There is no, so you can see games that are played far away (in effect, on another continent) without waiting time issues, but for games that you play, waiting time is (a) enjoy the game. To be, and (b) to be on an equal footing with other players with short latency connections, it must be short enough for the user). If the user is allowed to play, the app / game server 1521-1525, which provided the user with a "game finder" user interface, is suitable for the hosting service control server 401 to play a particular game. The configured app / game server 1521-1525 is started (ie, explored and started), the game is loaded from the RAID array 1511-1512, and then the hosting service control server 401 is currently hosting the game. The game is currently being played by instructing the inbound routing 1502 to transfer control signals from the user to the app / game game server that is in the game, as well as compressing the video / audio from the app / game server that hosted the "GameFinder" application. Requests the shared hardware compressor 1530 to switch to compressing video / audio from the hosting app / game server. "Game finder The vertical sync of the app / game service and the new app / game server hosting the game will not be synced, resulting in a probably time lag between the two syncs. The shared video compression hardware 1530 starts compressing video when the app / game server 1521-1525 completes the video frame, so the first frame from the new server is earlier than the full frame time of the old server. Completed in, this is before the compressed frame in front of it completes its transmission (for example, given the transmission time 992 in Figure 9b, the uncompressed frame 3 963 is completed half the frame time earlier. If so, it will affect the transmission time 992). In this situation, the shared video compression hardware 1530 ignores the first frame from the new server (for example, frame 4 964 is ignored (974)) and the client 415 is old. It holds the last frame from the server for a special frame time, and the shared video compression hardware 1530 begins compressing the next frame time video from the new app / game server that hosts the game. Visually, the transition from one app / game server to another is seamless. The hosting service control server 401 then notifies the app / game game server 1521-1525, which hosted the "game finder", to switch to idle until it is needed again. Client 415 holds the last frame from the old server for a special frame time (as well as 964 is ignored (974)), and the shared video compression hardware 1530 is a new app that hosts the game. Starts compressing the next frame-time video from the / game server. Visually, the transition from one app / game server to another is seamless. The hosting service control server 401 then notifies the app / game game server 1521-1525, which hosted the "game finder", to switch to idle until it is needed again.
0208The user can then play the game. And what's great is that the game is perceived to be played instantly (because it is loaded from the Raid array 1511-1512 to the app / game game server 1521-1525 at a speed of gigabit / sec), and the game. Is a strictly configured operating system for the game that has the ideal driver, registry configuration (for Windows), and no other applications running on the server that may conflict with the operation of the game. Loaded with a server that is exactly suitable for the game.
0209Also, as the user progresses through the game, each segment of the game is loaded from the RAID array 1511-1512 to the server at a rate of gigabit / sec (ie a 1 gigabyte load in 8 seconds), and the RAID array 1511-. Due to the huge storage capacity of the 1512 (which is very large and cost effective as it is a shared resource among many users), the geometric settings or other game segment settings are pre-computed. It can be stored in the RAID array 1511-1512 and loaded very quickly. Furthermore, since the hardware configuration and computing power of each app / game server 1521-1525 is known, pixels and vertex shaders can be pre-computed.
0210Therefore, the game starts almost instantaneously, runs in an ideal environment, and subsequent segments are loaded almost instantly.
0211However, beyond these effects, the user can see others playing the game (via the "game finder" and other means described above), and determine if they are interested in the game, and if so. If so, you can learn the secret from seeing others. The user can then instantly demo the game without having to wait for a large download and / or installation, and the user will probably play the game on a trial basis at a small cost or on a long-term basis. You can play instantly. The user can then play the game on a Windows PC, Macintosh, home TV receiver, or on a mobile phone while traveling, with a sufficiently short latency wireless connection. Also, all this can be done without having a physical copy of the game.
0212As mentioned above, users do not allow others to see their gameplay, allow their games to be seen after a delay, or allow their games to be seen by selected users. You can decide whether to allow or allow your game to be viewed by all users. Nevertheless, video / audio, in one embodiment, was stored in the delay buffer 1515 for 15 minutes, and could be done while the user was watching TV on a digital video recorder (DVR). Similarly, you can "rewind" and see your previous gameplay, pause, slow play, fast forward, and so on. In this example, the user plays the game, but the same "DVR" ability is available when the user uses the application. This is also useful for reviewing previous behavior and for other applications described below. Furthermore, if the game is designed to be rewound based on the use of game status information, such as changing the field of view of the camera, this "3D" The "DVR" ability is also supported, but it is required to design the game to support it. The "DVR" ability using the delay buffer 1515, of course, occurs when the game or application is used. For games that work with the game or application only for video, but with 3D DVR capabilities, the user can control the "fly through" in 3D of previously played segments. , And the resulting video can be recorded in the delay buffer 1515 to record the game state of the game segment. Thus, a particular "flythrough" is recorded as compressed video, but the game state is also recorded, so different flythroughs may be considered for the same segment of the game at a later date.
0213As described below, each user in the hosting service 210 has a "user page" on which information and other data about them can be posted. Among the things users can post are video segments from gameplay they have saved. For example, if a user overcomes a particularly difficult challenge in a game, the user "rewinds" to just before the point where he or she has achieved great results in the game, and then some time for another user to see. The hosting service 210 can be instructed to save a width (eg, 30 seconds) video segment to the user's "user page". To achieve this, the user simply plays the video stored in the delay buffer 1515 to the RAID array 1511-1512 and uses it to index that video segment on the user's "user page". This is a problem with the game server 1521-1525.
0214When the game has the above-mentioned 3D DVR capability, the game state information required for the 3D DVR can also be recorded by the user and used for the user's "user page".
0215A "Game Finder" application if the game is designed to have an "audience" (ie, a user who can navigate the 3D world and observe it without involving any action) in addition to the active player. Allows users to join the game as spectators and players. From a realization point of view, there is no difference to the hosting system 210 even if the user is an spectator rather than an active player. The game is loaded on the app / game server 1521-1525, and the user controls the game (eg, controls a virtual camera looking at the world). The only difference is the user's gaming experience.
0216<u style="single">Cooperation of multiple users</u> Another feature of hosting service 210 is the ability of multiple users to collaborate while watching live video, even when using very different viewing devices. This is useful both when playing games and when using applications.
0217Many PCs and mobile phones are equipped with video cameras and have the ability to perform real-time video compression, especially when the video is small. A small camera that can be attached to a television is also available, and real-time compression can be embodied in software or using one of many hardware compression devices for compressing video. , Not difficult. Also, many PCs and all mobile phones have a microphone, and a headset can be used with the microphone.
0218When such a camera and / or microphone is combined with local video / audio compression capabilities (particularly using the short latency video compression techniques described herein), the user can view the video from the user's home 211. And / or allow audio to be sent to the hosting service 210 along with the control data of the input device. When such techniques are used, the capabilities shown in FIG. 19 can be achieved, i.e., users can have their video and audio 1900 appear on screens in another user's game or application. An example of this is a multiplayer game in which teammates collaborate in car racing. The user's video / audio can be selectively viewed / listened to only by teammates. And since there is virtually no waiting time, using the techniques described above, players can talk and move with each other in real time without noticeable delays.
0219This video / audio integration is achieved by arriving compressed video and / or audio from the user's camera / microphone as inbound internet traffic 1501. Inbound routing 1502 then routes the video and / or audio to the app / game game server 1521-1525, which is allowed to watch / listen to it. The user of each app / game game server 1521-1525, who then chooses to use video and / or audio, unzips it and, as indicated in 1900, within the game or application. Integrate so that it appears in.
0220The example in Figure 19 shows how such cooperation is used in games, but such cooperation is a very powerful tool for applications. In New York City, a larger building was designed by a Chicago architect for a New York-based real estate developer, but the decision included a financial investor who happened to be at a Miami airport while traveling, and Consider a situation in which some design factors of a building need to be judged on how to harmonize with the surrounding building in order to satisfy both investors and real estate developers. Construction companies have high-definition monitors with cameras mounted on PCs in Chicago, real estate developers have laptops with cameras in New York, and investors have mobile phones with cameras in Miami. doing. Construction companies can use hosting service 210 to host powerful construction design applications that can perform highly realistic 3D rendering of a large database of buildings in New York City and of buildings under design. You can use the database. Construction design applications run on one of the app / game servers 1521-1525, or some of them if they require a great deal of computational power. Each of the three users in different locations is connected to the hosting service 210, and each has a simultaneous view of the video output of the construction design application, which is the given equipment and network that each user has. Shared hardware compression 1530 makes it suitable for connectivity characteristics (for example, construction companies can see a 2560x1440 60fps display through a 20Mbps commercial internet connection, and New York real estate developers can see their own. You can watch 1280x720 60fps video over a laptop's 6Mbps DSL connection, and investors can see 250Kbps cellular data on their mobile phones. You can watch 320x180 60fps video via the connection. Each party listens to the other party's voice (conference calls are handled by one of the many conference call software packages widely available on the app / game server 1521-1525), and the operation of buttons on the user input device. Through, users can make their own video appearances using local cameras. As the meeting progresses, the architect can show in a very photorealistic 3D rendering what the building will look like when it is rotated and flying near other buildings in the area, and The same video can be viewed by all parties at the resolution of each party's display device. It's not a problem that none of the local devices used by the parties can handle 3D animations with such realism, but will they download the huge database required to render the buildings around New York City? Or it goes without saying that it is remembered. From each user's point of view, the user simply has an incredible realism and seamless experience, despite the distance and the different local devices. And when one party wants to be able to see his face and convey his emotions well, he can do so. In addition, if either a real estate developer or an investor wants to gain control of a construction program and use their own input device (keyboard, mouse, keypad or touch screen), there is no perceived waiting time. It can and will respond (assuming those network connections do not have unreasonable latency). For example, in the case of a mobile phone, the waiting time is very short when the mobile phone is connected to the airport's WiFi network. However, when using the cellular data networks available today in the United States, you will probably suffer from significant delays. In addition, most of the meetings that investors are watching
0221Finally, at the end of the co-conference call, real estate developers and investors issue and sign their comments from the hosting service, and the construction company "rewinds" the conference video recorded in the delay buffer 1515. , And the comments, facial expressions, and / or actions added to the 3D model of the building created during the meeting can be reviewed. If there are specific segments that you want to save, you can move those segments of video / audio from the delay buffer 1515 to the RAID array 1511-1512 for record storage and later playback.
0222Also, from a cost perspective, if the architect only needs to use computational power and a large database in New York City for a 15-minute conference call, he owns a high-power workstation and an expensive copy of the large database. Instead of buying, you only have to pay for the time you spend the resource.
0223<u style="single">Video-rich community service</u> Hosting service 210 enables an unprecedented opportunity to establish a video-rich community service on the Internet. FIG. 20 shows a normative "user page" for game players in hosting service 210. Like the "Game Finder" application, the "User Page" is an application that runs on one of the app / game servers 1521-1525. All thumbnails and video windows on this page show constantly moving video (loops if the segment is short).
0224By using a video camera or uploading a video, the user (whose username is "KILLHAZARD") can post their own Video 2000, which other users can see. .. This video is stored in the RAID array 1511-1512. Also, when another user arrives at KILLHAZARD's "user page", if KILLHAZARD is using hosting service 210 at that time (the user sees his "user page" to see him). No matter what he does (assuming he forgives), the live video 2001 is shown. This is achieved by the app / game server 1521-1525 hosting the "user page" application requested by server control system 401, whether KILLHAZARD is active, and if so, the app he uses. Achieved by / game server 1521-1525. A compressed video stream of the appropriate resolution and format is then sent to the app / game server 1521-1525 running the "User Page" application, using the same method used by the "Game Finder" application. And it is displayed. If the user selects a window in KILL HAZARD's live gameplay and then clicks on its input device appropriately, the window will be zoomed in (again using the same method as in the "Game Finder" application). The live video then fills the screen with the resolution of the viewing user's display device 422, which is suitable for the characteristics of the viewing user's Internet connection.
0225An important effect of this solution over the past is that users viewing the "user page" can see the live play of the game they do not own, and, very goodly, play the game. You don't have to have a local computer or game console that you can. This provides a good opportunity for the user to see the user shown on the "user page" which is "in action" to play the game, and also an opportunity for the viewing user to learn about the game they want to try or get better at.
0226Video clips recorded or uploaded by KILLHAZARD's buddy 2002 are also shown on the "User Page", and below each video clip is text indicating whether the buddy is playing the game online (eg,). six_shot is playing the game "Eragon", MrSnuggles99 is offline, etc.). By clipping a menu item (not shown), the buddy's video clip shows what the recorded or uploaded video is, what the buddy currently playing the game on hosting service 210 then does in the game. Switch to a live video of what you're doing. Therefore, this is a game finder group for buddies. If a buddy's game is selected and the user clicks on it, the full screen is zoomed in and the user can see the game played live on the full screen.
0227Again, the user watching the buddy's game does not own a copy of the game or local compute / game console resources to play the game. Watching the game is practically instantaneous.
0228As mentioned above, when a user plays a game on the hosting service 210, the user "rewinds" the game, finds the video segment he wants to save, and then saves the video segment to his "user page". These are referred to as "Brag Clips". Video Segment 2003 is all "Proud Clips" 2003 saved by KILL HAZARD from previous games played. The number 2004 indicates how many times the "pride clip" has been seen, and when the "pride clip" is seen, the user has the opportunity to rate it and the orange keyhole-shaped icon 2005. The number indicates how high the rating is. "Pride Clip" 2003 constantly loops with the rest of the video on a "user page" when the user views it. If the user selects and clicks on one of the "Pride Clips" 2003, it will zoom in and the "Pride Clip" 2003 will play, pause, rewind, fast forward, step through, etc. Presented with a DVR control that allows you to do.
0229Playback of "Pride Clip" 2003 shows the compressed video segment stored in the RAID array 1511-1512 when the user records the "Pride Clip", decompresses it, and plays it. It is embodied by loading the 1525.
0230The "Proud Clip" 2003 is also a "3D DVR" video segment (ie, a game state sequence from a game that can be played and allows the user to change the viewpoint of the camera), which is such an ability. Is from a game that supports. In this case, game state information is stored in addition to the specific "fly-through" compressed video recording created by the user when the game segment was recorded. When the "User Page" is viewed and all thumbnails and video windows are constantly looping, the 3D DVR "Pride Clip" 2003 records as compressed video when the user records a "flythrough" of the game segment. The "Pride of Clips" 2003 that was made is constantly looped. However, when the user selects the 3D DVR "Pride Clip" 2003 and clips it, in addition to the DVR control that allows the compressed video "Pride Clip" to be played, the user can use the 3D for the game segment. You can click the button that gives you the DVR ability. They can control the camera's "fly-through" during the game segment on their own, and if they want (and the user who owns the user page allows it), another "pride clip" "Fly-through" can be recorded in compressed video format, which will then be available to other viewers of the user page (immediately or by the user page owner with a "pride clip". After having the opportunity to revisit).
0231This 3D DVR "Pride Clip" 2003 capability is enabled by running a game that is trying to play the recorded game state information on another app / game server 1521-1525. Since the game can be run almost instantaneously (as mentioned above), the game is run with limited play to the game state recorded by the "pride clip", and then the user delays the compressed video buffer 1515. It's not difficult to be able to perform a "fly-through" on the camera while recording to. The game is deactivated when the user completes the "flythrough" run.
0232From the user's point of view, activating the "flythrough" on the 3D DVR "Pride Clip" 2003 requires less effort than controlling the linear "Pride Clip" 2003 DVR control. They need not know anything about the game or how to play the game. They are just virtual camera operators staring at the 3D world during a game segment recorded by another.
0233Users can also overdub their own audio into "proud clips" that are recorded or uploaded from the microphone. In this way, "pride clips" can be used to generate custom animations using characters and actions from the game. This animation technique is commonly known as "machinima".
0234Achieve different skill levels as the user progresses through the game. The games played report their achievements to the service control system 401, and their skill levels are shown on the "user page".
0235<u style="single">Two-way animated advertising</u> Online advertising has moved from text to still video, to video, and now to interactive segments that are typically embodied using animated thin clients like Adobe Flash. The reason for using animated thin clients is that users typically cannot tolerate delays in the perks that products and services are recommended to them. Also, thin clients run on very low performance PCs, so advertisers can have a high degree of confidence that bidirectional ads will work properly. Unfortunately, animation thin clients like Adobe Flash have limited interactivity and limited experience (to reduce download time).
0236FIG. 21 shows a two-way advertisement in which the user selects the exterior and interior colors of the car while rotating the car in the showroom and showing in real-time ray tracing what the car looks like. The user selects the avatar to drive the car, and then the user can retrieve the car to drive on a race track or in an exotic location such as Monaco. The user can choose a larger engine or better tires and then see how the modified configuration affects the acceleration performance of the car or the bite of the road surface.
0237Of course, advertising is, in effect, an elaborate 3D video game. However, for ads that can be played on a PC or video game console, it will probably require a 100MB download, and in the case of a PC, it will require the installation of special drivers and the PC will have sufficient CPU or GPU computing power. If it is missing, it may not be executed at all. Therefore, such an advertisement is not possible with the conventional configuration.
0238With hosting service 210, such advertisements launch almost instantly, and are fully executed, whatever the capabilities of the user's client 415. Therefore, they launch faster than thin client bidirectional ads, are significantly more experienced, and are very reliable.
0239<u style="single">Streaming geometry during real-time animation</u> RAID array 1511-1512 and inbound routing 1502 provide RAID array 1511-1512 to ensure on-the-fly delivery of geometry in the middle of gameplay or applications during real-time animation (eg, flythrough with complex databases). And fast and short latency data rates can be provided so that video games and applications that rely on inbound routing 1502 can be designed.
0240In traditional systems such as the video game system shown in Figure 1, mass storage devices that can be used, especially in real home appliances, have geometric shapes during gameplay, except in situations where the required geometry can be predicted to some extent. Is much slower to stream. For example, in a drive game where a particular road is present, the geometry of the building coming into view is reasonably predictable, and mass storage is where the approaching geometry is located. You can search for a place in advance.
0241However, in complex scenes with unpredictable changes (for example, in combat scenes with complex characters all over), the RAM of the PC or video game system is completely geometrically shaped to the object currently in view. If it is buried and then the user suddenly redirects those characters to the field of view behind them, there will be a delay before the geometry can be displayed if it is not preloaded in RAM. It will be.
0242In Hosting Service 210, the RAID array 1511-1512 can stream data above Gigabit Ethernet speeds, and in SAN networks achieve 10 Gigabit / s speeds over 10 Gigabit Ethernet or other networks. Can be done. 10 Gigabit / sec loads gigabytes of data in less than a second. At 60fps frame time (16.67ms), it can load almost 170 megabits (21MB) of data. Of course, rotating media in a RAID configuration still incurs latency greater than one frame time, but flash-based RAID storage devices are, after all, as large as rotating media RAID arrays, and such You don't have to wait long. In one embodiment, high volume RAM write-through caching is used to provide access with very short latency.
0243Therefore, with a sufficiently high network speed and a sufficiently short latency mass storage, the geometry can be transferred to the app / game game server 1521-1525 as fast as the CPU and / or GPU can process 3D data. Can be streamed. Thus, in the above example where the user suddenly turns the character and looks back, the geometry for all the characters behind can be loaded before the character completes the rotation, and thus for the user. It feels like you're in a geometric world as real as a live action.
0244As mentioned above, one of the last untapped things in photorealistic computer animation is the human face, which is sensitive to imperfections by the human eye, resulting in some errors from the photorealistic face. It may cause a negative reaction from the viewer. Figure 22 shows Contour<sup>TM</sup>"Reality Capture Technology" (a simultaneously pending patent application, each transferred to the transferee of this CIP application, i.e., No. 10 / 942,609 filed on September 15, 2004, "Apparatus and method for capturing the motion of a" "Performer", No. 10 / 942,413 "Apparatus and method for capturing the expression of a performer" filed on September 15, 2004, No. 11 / 066,954 "Apparatus and method" filed on February 25, 2005. for improving marker identification within a motion capture system , filed on March 10, 2005, No. 11 / 077,628 Apparatus and method for performing motion capture using shutter "synchronization", No. 11 / 255,854 filed on October 20, 2005, "Apparatus and method for performing motion capture using a random pattern on capture surfaces", No. 11 / 449,131 filed on June 7, 2006. System and method for performing motion capture using phosphor application techniques, No. 11 / 449,043 System and method for performing motion capture by strobing a fluorescent lamp, filed June 7, 2006, June 7, 2006. No. 11 / 449,127 System and method for three dimensional capture of stop-motion animated filed in Raw performances captured using characters ) have a very smooth captured surface, followed by a high polygon number tracking surface (ie, polygonal movements exactly follow facial movements). Finally, a video of the live performance is mapped to the tracking surface to give a photorealistic result when forming a textured surface.
0245Current GPU technology can render the number and texture of polygons on the tracking surface and illuminate the surface in real time, but the polygons and textures change from frame time to frame time (which gives the most photorealistic results). If it does), it will rapidly consume all available RAM on a modern PC or video game console.
0246Using the streaming geometry technology described above, it continuously feeds the geometry to the app / game game server 1521-1525, continuously animating the photorealistic face and making it almost distinguishable from the raw moving face. It is practical to be able to create a video game with a dull face.
0247<u style="single">Integration of linear content with bidirectional features</u> Video, television programming and audio material (collectively "linear content") are widely available to home and office users in many forms. Linear content can be acquired on physical media such as CD, DVD, HD-DVD and Blu-ray media. It can also be recorded by DVR from satellite and cable TV broadcasts. It can also be used as pay-per-view (PPV) content through satellite and cable TV, and as video-on-demand (VOD) on cable TV.
0248More and more linear content is now available as downloaded content and streaming content through the Internet. Today, there is really more than one place to experience all the features associated with linear media. For example, DVDs and other video optical media typically have bidirectional features that are not available elsewhere, such as director commentary, "making of" features, and so on. Online music sites have cover technology and song information that are generally not available on CDs, but not all CDs are available online. Also, websites associated with television programs often have special features, blogs, and sometimes comments from actors or creative staff.
0249In addition, many video or sporting events are often released with linear media (in the case of video) or (in the case of sporting) videos closely linked to real-world events (eg, player trades). There are often games.
0250Hosting service 210 is well suited for linking and distributing linear content with different forms of related content. Indeed, video distribution is no longer the challenge of delivering highly interactive video games, and hosting service 210 can deliver linear content to a wide range of home or office devices or mobile devices. FIG. 23 shows the normative user interface page of hosting service 210 showing the selection of linear content.
0251However, unlike most linear content delivery systems, the hosting service 210 provides related bidirectional components (eg, menus and features on DVD, bidirectional overlays on HD-DVD, and Adobe Flash animations on websites (discussed below). )) Can also be delivered. Therefore, the limitation of client device 415 no longer introduces a limitation on which features are available.
0252In addition, the hosting service 210 can dynamically and in real time link the linear content with the video game content. For example, if a user sees a Quidditch match in a Harry Potter movie and wants to try playing Quidditch, they can pause the movie and immediately move to the Quidditch segment of the Harry Potter video game with the click of a button. it can. After playing a Quidditch match, another click on the button will instantly restart the movie.
0253With photorealistic graphics and production techniques that make photographicly captured videos indistinguishable from live motion characters, users can move from quidditch games in live motion movies to video games in the hosting services described here. When moving to a quidditch game, the two scenes are virtually indistinguishable. This gives directors of both linear and interactive (eg video game) content a whole new creative option as the line between the two worlds is indistinguishable.
0254The hosting service architecture shown in Figure 14 can be used to give viewers control of a virtual camera in a 3D movie. For example, in a scene that occurs inside a train vehicle, the viewer can control a virtual camera to look around the vehicle as the story progresses. It assumes that all 3D objects (assets) in the vehicle are available, enough computing power to render the scene in real time, and the original movie.
0255And even non-computer-generated entertainment can provide very exciting bidirectional features. For example, the 2005 video "Pride and Prejudice" has many scenes in a splendid old English mansion. For some scenes in the mansion, the user can pause the video and then control the camera to shoot the mansion tower, or perhaps the surrounding area. To do this, you can carry the camera through the mansion with a fisheye lens so you don't lose track of your position, just as Apple's traditional QuickTime VR was. The various frames are then converted so that the video is stored in the RAID array 1511-1512 along with the movie without distortion and can be played back when the user chooses to go to the virtual tower.
0256Sporting events allow users to stream live sporting events, such as basketball games, through hosting service 210 for viewing on regular TV. After the user sees a particular play, the player can start the video game of the game (eventually with a basket player that looks as photorealistic as the real player) in the same position, and the user. (Probably each with one player's control) can replay and see if it works better than the player.
0257The hosting service 210 described here is very well suited to support this futuristic world. This is because it can hold computational power and large storage resources that are not possible to install in home or most office settings, and in home settings it will always have older generations of PCs and video games. That's because its computing resources are the latest computing hardware available and always up-to-date. And in hosting service 210, all the complexity of this calculation is hidden from the user, so it is as easy as switching TV channels from the user's point of view, even with a very sophisticated system. In addition, the user has access to all computational power and experience where computational power is obtained from client 415.
0258<u style="single">Multiplayer game</u> A server that can communicate with the app / game game server 1521-1525 through the network of inbound routing 1502 in that the game is a multiplayer game, and does not run on the hosting service 210 on a network bridge to the internet (not shown) Can communicate with game machines. When playing multiplayer games on a typical Internet computer, the app / game game server 1521-1525 says it has very fast access to the Internet (compared to when the game is run on a home server). Although it has advantages, it is limited by the ability of other computers to play games on slow connections, and the game server on the Internet is the minimum public denominator, which is a home computer on a relatively slow consumer Internet connection. It is also potentially limited by being designed to accept.
0259But when multiplayer games are played entirely within the server center of hosting service 210, a world of difference can be achieved. Each app / game game server 1521-1525 hosting a game for the user, with other app / game game servers 1521-1525, has a very fast, very short latency connection and a huge amount of very fast storage. It is interconnected with a server that hosts centralized control for multiplayer games with arrays. For example, if Gigabit Ethernet is used for the network of inbound routing 1502, the app / game game server 1521-1525 will communicate with each other and with a latency of potentially less than 1ms at gigabit / s speeds. It also communicates with the server that hosts centralized control for multiplayer games. In addition, the RAID array 1511-1512 responds very quickly and can transfer data at gigabit / s speeds. For example, in a traditional system restricted to game clients running on a PC or game console at home, the user customizes the character in terms of appearance and clothing so that the character has a large number of geometric shapes and the character's unique behavior. If the character enters another user's field of view, the user waits until the long, slow download is complete and all geometry and behavior data is loaded into the computer. You will have to. Within the hosting service 210, the same download can be done at gigabit / s speeds from the RAID array 1511-1512 via the corresponding Gigabit Ethernet. Gigabit Ethernet is 100 times faster, even if home users have an 8Mbps Internet connection, which is very fast by today's standards. Therefore, what takes one minute for a high-speed Internet connection is less than one second for Gigabit Ethernet.
0260<u style="single">Top player groups and tournaments</u> Hosting service 210 is very well suited for tournaments. The local client doesn't run the game, so users don't have a chance to cheat. Also, since egress routing 1540 can multicast UDP streams, hosting service 210 can broadcast major tournaments to thousands of people in the spectator at once.
0261In fact, when there are several video streams that have become so popular that thousands of users receive the same stream (eg, showing a view of a major tournament), Akamai or Akamai or for mass distribution to many client devices 415. It is more efficient to send the video stream to a "content delivery network" (CDN) like Limelight.
0262Similar efficiency levels can be obtained when using a CDN to show the "Game Finder" page of a top player group.
0263For major tournaments, you can use a live, well-known announcer to broadcast live during a match. Many users are watching major tournaments, but relatively few are playing in tournaments. Voices from well-known announcers can be routed to the app / game game server 1521-1525, which hosts the user playing in the tournament and also hosts the audience mode copy of the game in the tournament, and the voice is the game voice. Can be overdubbed on top. The footage of a well-known announcer can also be overlaid on the game, and perhaps on the audience's view.
0264<u style="single">Accelerate web page loading</u> The Worldwide Web, its primary transport protocol, Hypertext Transfer Protocol (HTTP), was conceived and defined in an era when only businesses had high-speed Internet connections and consumers online were using dial-up modems or ISDNs. It was done. At that time, the "gold standard" for high-speed connections was the T1 line, which gives a data rate of 1.5 Mbps symmetrically (ie, at data rates equal to both directions).
0265Today, the situation is completely different. In much of the developed world, the average home connection speed through a DSL or cable modem connection has a much higher downstream data rate than the T1 line. In fact, in some parts of the world, fiber-to-the-curb carries data rates as high as 50 to 100 Mbps to the home.
0266Unfortunately, HTTP has not been configured (or even materialized) to effectively take advantage of these rapid speed improvements. A website is a collection of files on a remote server. Very simply, HTTP requests the first file, waits for the file to be downloaded, then requests the second file, waits for the file to be downloaded, and And so on. In fact, HTTP allows more than one "open connection", i.e. allows you to request more than one file at a time, but a request to prevent the agreed standard (and the web server from overloading). ) Allows very few open connections. Moreover, because of the way web pages are constructed, browsers are often unaware of multiple simultaneous pages available for immediate download (ie, only after parsing the page, as well as new video. It becomes clear that the file needs to be downloaded). Therefore, website files are essentially loaded one by one. And because of the request and response protocols used by HTTP, there is a latency of approximately 100ms associated with each file loaded (when accessing a typical web server in the United States).
0267For relatively slow connections, this does not introduce a major problem. This is because the download time of the file itself dominates the latency of the web page. But as connection speeds increase, problems begin to occur, especially on complex web pages.
0268The example shown in Figure 24 shows a typical commercial website (this particular website is from a brand of shoes for top athletes). This website has 54 files. These files include HTML, CSS, JPEG, PHP, Java® Script, and Flash files, as well as video content. A total of 1.5 Mbytes must be loaded before the page is raw (ie, the user clicks on it and starts using it). Many files have many reasons. One is a complex and sophisticated web page, and another is dynamically dynamic with information about the user accessing the page (eg, the user's home country, language, whether they have previously purchased, etc.). Web pages that are assembled based on, and different files are downloaded based on all these factors. Still, this is a very typical commercial web page.
0269Figure 24 shows the amount of time it takes for a web page to grow as the connection speed increases. At a 1.5Mbps connection speed of 2401, using a traditional web server with a traditional web browser takes 13.5 seconds for a web page to grow. At a 12Mbps connection speed of 2402, the load time is reduced to 6.5 seconds, or about twice as fast. However, with a 96Mbps connection speed of 2403, the load time is only reduced to about 5.5 seconds. The reason is that at such a high download speed, the time to download the file itself is minimal, but there is still a latency per file, about 100ms each, resulting in 54 files * 100ms = 5.4. There is a waiting time of seconds. So, no matter how fast you connect to your home, this website always takes at least 5.4 seconds for it to come to life. Another factor is server-side queuing, that is, each HTTP request is added to the back of the queue, so for busy servers this has a big impact. That's because every time you get a small item from a web server, the HTTP request has to wait for that change of direction.
0270One way to solve these problems is to drop or redefine HTTP. Or perhaps the website owner successfully merges the file into a single file (eg in Adobe Flash format). But as a practical matter, the company and many others have invested heavily in its website architecture. In addition, some homes have a connection of 12-100Mbps, but the majority of homes are still at slow speeds, and HTTP also works at low speeds.
0271One alternative is to host the web browser for the app / game server 1521-1525 and host the files for the web server for the RAID array 1511-1512 (or potentially host the web browser app / In the RAM or local storage of game server 1521-1525). Due to the very fast interconnection through inbound routing 1502 (or to local storage), there is a minimum wait per file using HTTP rather than a 100ms wait per file using HTTP. It will be time. Therefore, instead of the user at home accessing the web page via HTTP, the user can access the web page through the client 415. Then, on a 1.5Mbps connection (because this web page doesn't require significant bandwidth for its video), the web page grows in less than 1 second per 2400 lines. In essence, there is no waiting time for the web browser running on the app / game server 1521-1525 to display the raw page, and detectable waiting for the client 415 to display the video output from the web browser. no time. When the user mouses and / or types on a web page, the user's input information is sent to a web browser running on the app / game server 1521-1525, and the web browser responds accordingly.
0272One drawback to this solution is that if the compressor constantly sends video data, bandwidth will be used even if the web page is static. This can be ameliorated by configuring the compressor to send data only when the web page changes (and only then) and then only to the portion of the changed page. There are some web pages with constantly changing flushing banners, etc., but such web pages tend to be annoying, and web pages are usually quiet unless there is a reason for something to move (eg, a video clip). It will be a typical one. For such web pages, perhaps less data is sent using the hosting service 210 than traditional web servers. This is because only the video that is actually displayed is transmitted, and neither the thin client executable code nor large objects that are never seen, such as rollover video, are transmitted.
0273Therefore, using hosting service 210 to host a legacy web page, the load time of the web page can be reduced to the point that opening the web page is equivalent to switching TV channels, ie. Web pages are virtually instantly raw.
0274<u style="single">Easier debugging of games and applications</u> As mentioned above, video games and applications with real-time graphics are very complex applications and may typically contain bugs when they are put on the market. Software developers have some way to get user feedback about bugs and reduce machine state after a crash, but strictly identify what happened to the game or real-time application that caused it to crash or malfunction. It's very difficult to do.
0275When the game or application runs on the hosting service 210, the video / audio output of the game or application is constantly recorded in the delay buffer 1515. In addition, the watchdog process regularly reports to the hosting service control system 401 that each app / game server 1521-1525 is running and that the app / game server 1521-1525 is running smoothly. If the watchdog process fails to report, the server control system 401 will attempt to communicate with the app / game server 1521-1525 and, if successful, will collect it no matter what machine state is obtained. .. Whatever information is obtained, it is sent to the software developer along with the video / audio recorded in the delay buffer 1515.
0276Therefore, when a game or application software developer gets a crash notification from hosting service 210, he gets a frame-by-frame record of what caused the crash. This information is of great value in tracking down and fixing bugs.
0277Also, if the app / game server 1521-1525 crashes, the server will be restarted at the most recent restartable point, and the user will be given a message apologizing for the technical difficulties.
0278<u style="single">Resource sharing and cost savings</u> The systems shown in Figures 4a and 4b offer a variety of benefits to both end users and game and application developers. For example, home and office client systems (eg, PCs or game consoles) are typically used for only a short time a week. Nielsen Entertainment Active Gamer Benchmark According to a press release on October 5, 2006 by EDATE =), aggressive gamers spend an average of 14 hours a week playing video game consoles, and handheld about 17 hours a week. There is. The report also states that aggressive gamers spend an average of 13 hours a week on all gameplay activities (including console, handheld and PC gameplay). Given the high number of console video game play times, a week is 24 * 7 = 168 hours, which is the amount of time a video game console is in a week at an aggressive gamer's home. It means that only 17/168 = 10% of them are used. That is, the video game console is idle for 90% of the time. It's a high-cost video game console, and if the manufacturer is subsidizing such a device, it's a very inefficient use of expensive resources. PCs in the office are also typically used for only a short time of the week, especially Autodesk. This is the case with non-portable desktop PCs, which are often required for high-end applications like Maya. Some companies are open all the time and holidays, and some PCs (eg portables carried to work at home at night) are used all the time and holidays, but most of the company's activities. Tends to be centered around 9am to 5pm in the designated business hours zone from Monday to Friday, rarely on holidays and rest times (such as at lunch), and most PCs used The use of desktop PCs tends to follow these business hours, as it occurs while the user is actively involved in the PC. Assuming that the PC is used constantly from 9 AM to 5 PM, 5 days a week, it means that the PC is used 40/168 = 24% of the time of the week. High-performance desktop PCs are a very expensive investment for a company, which reflects a very low level of utilization. Schools that teach on desktop computers use computers even in smaller hours of the week and vary based on class hours, but most classes are held during the daytime hours, Monday through Friday. Therefore, in general, PCs and video game consoles are only used for a small portion of the weekly hours.
0279In particular, many people are in the office or school during the daytime hours from Monday to Friday on weekdays, so these people generally do not play video games during these hours, and therefore play video games. Playing is generally during other times such as nights, weekends and holidays.
0280Given the hosting service configuration shown in Figure 4a, the usage patterns described in the two paragraphs above result in very efficient use of resources. Obviously, a user who can be serviced by hosting service 210 at a given time, especially if the user requires a real-time response to a complex application such as a sophisticated 3D video game. There is a limit to the number. However, unlike home video game consoles and office PCs, which are typically idle most of the time, the server 402 can be reused by different users at different times. For example, a high-performance server 402 with high-performance dual CPUs and dual GPUs and a large amount of RAM can be used by companies and schools from 9 am to 5 pm on weekdays, but on nights, weekends and holidays. Available to gamers who play sophisticated video games. Similarly, low-performance applications are Celeron Available by companies and schools during company time on low performance server 402 with CPU, no GPU (or very low end GPU) and limited RAM, and low performance games Can take advantage of the low-performance server 402 when it is not during company hours.
0281Further, in the hosting service configuration described here, resources are efficiently shared among thousands of users, if not millions. In general, online services make up only a small percentage of the total user base that uses the service at a given time. Given Nielsen's video game usage statistics mentioned above, it's easy to see why. If aggressive gamers play console games for only 17 hours a week, the peak game usage will be nights (5-AM12, 7 * 5 days = 35 hours / week) and weekends (AM8-AM12, 16). Assuming during typical non-working non-company hours (* 2 = 32 hours / week), there are 35 + 32 = 65 peak hours per week for 17 hours of gameplay. The exact peak user load on the system is difficult to estimate for many reasons: some users play during off-peak hours, there is a collective peak of users during some daytime hours, and peak hours play. Influenced by the format of the game (for example, children's games are probably played early in the night), and so on. However, if the average number of hours a gamer plays is probably much less than the number of hours a gamer plays a game on a day, only a portion of the hosting service 210's users will use it at a given time. is there. For this analysis, it must be assumed that the peak load is 12.5%. Therefore, only 12.5% of the computation, compression and bandwidth resources are used at a given time, and as a result, a given user is asked to play a game of a given performance level by reusing the resources. Only 12.5% of the hardware cost to support.
0282Moreover, if one game and application requires more computing power than another, resources can be dynamically allocated based on the game played by the user or the application being executed. Therefore, a user who chooses a low performance game or application is assigned a low performance (cheap) server 402, and a user who chooses a high performance game or application is assigned a high performance (more expensive) server 402. Be done. In fact, a given game or application has a low performance and high performance category of the game or application, and the user retains the user's behavior on the lowest cost server 402 that meets the needs of the game or application. As such, between game or application compartments, one server 402 can be switched to another. The RAID array 405, which is much faster than a single disk, can also be used with the low performance server 402, benefiting from high disk transfer rates. Therefore, the average cost per server 402 across all games played or applications used is significantly lower than the cost of the most expensive server 402 playing the highest performance games or applications, and the low performance server 402 is Derived disk performance benefits from RAID array 405.
0283Furthermore, the server 402 of the hosting service 210 is nothing more than a PC motherboard with no disk or peripheral interface other than the network interface, and will eventually be integrated into a single chip with only a high speed network interface to the SAN 403. .. Also, the RAID array 405 is probably shared by a much larger number of users than a disk, so the disk cost per active user is much lower than a single disk drive. All of this equipment is probably in a rack in an environmentally controlled server room environment. If server 402 fails, it can be easily repaired or replaced by hosting service 210. In contrast, a PC or gaming console in a home or office must be able to survive moderate wear and tear from being hit or dropped, requires a housing, and has at least one disk drive. Must survive in adverse environmental conditions (eg, stuffed into an overhead AV cabinet with other gear), require a service guarantee, must be packaged and shipped, and probably sold by a retailer that collects a retail margin. Must be a robust stand-alone device to be used. In addition, PCs or game consoles are expected to be the most computationally powerful games or applications that will be used at some point in the future, even if low-performance games or applications (or game or application categories) are played most of the time. It must be configured to satisfy the peak performance of. And if a PC or console fails, repairing it is an expensive and time-consuming process (which adversely affects manufacturers, users and software developers).
0284Therefore, if the system shown in Figure 4a gives users an experience comparable to local computing resources, then the architecture shown in Figure 4a allows home, office or school users to experience a given level of computing power. It becomes quite cheap to give its computing power through.
0285<u style="single">Eliminating the need for upgrades</u> Moreover, users no longer have to worry about upgrading their PCs and / or consoles to play new games or handle new applications with higher performance. The games or applications in the hosting service 210 are available to the user regardless of what form of server 402 is required for the games or applications, and all games and applications are available almost immediately (ie, ie). Server 402, which runs a given game or application, is loaded quickly from the RAID array 405 or local storage on server 402 and is properly performed with the latest updates and bug fixes (ie, the software developer runs a given game or application). You can choose the server configuration that is ideal for you, then you can configure the server 402 with the best drivers, and over time, the developer updates to all copies of the game or application in the hosting service 210. Bug repairs, etc. can be given at once). In fact, after the user started using the hosting service 210, the user probably found a game and application that would continue to give a great experience (eg, through updates and / or bug fixes), and the user was a year ago. Discovered a year later that a new game or application would be available for a service 210 that uses computing technology (eg, a high-performance GPU) that did not exist in, and thus play or run the game a year later. You will discover that it was impossible for users to buy the technology a year ago. The computational resources that play the game or run the application are invisible to the user (ie, from the user's point of view, the user starts running the game or application almost instantly, similar to switching TV channels. (Because you simply select), the user's hardware is "upgraded" even if the user is constantly unaware of the upgrade.
0286<u style="single">Eliminating the need for backup</u> Another major issue for users in businesses, schools and homes is backup. Information stored on the local PC or video game console (eg, in the case of a console, the user's game performance and ranking) is lost if the disc fails or is accidentally erased. Many applications that provide manual or automatic backups for PCs are available and can upload the state of the game console to an online server for backup, but local backups are typically another local disk (or other non-volatile). It must be copied to a sexual storage device) and stored and organized somewhere safe, and backups for online services are slowly upstream available through a typical low cost internet connection. Because of its speed, it is often limited. In the hosting service 210 of FIG. 4a, the data stored in the RAID array 405 is not lost even if the disk fails using the conventional RAID configuration technology well known to those skilled in the art. The server center expert containing the disk can be notified to replace the disk and configure the RAID array to automatically update to fail again. In addition, since all disk drives are close to each other, and there is a high-speed local network between them through SAN402, all disk systems are regularly backed up by secondary storage in the server center, and this secondary storage. It is not difficult to store the device in a server center or move it to a remote location. From the user's point of view of hosting service 210, the data is simply always secure and you don't have to think about backups.
0287<u style="single">Access to the demo</u> Users often want to try it before purchasing a game or application. As mentioned above, there are traditional means of demonstrating games and applications (the verb form of "demo" means trying the demonstration version, which is also referred to as the noun "demo"). Each of them suffers from restrictions and / or inconveniences. Hosting service 210 makes it easy and convenient for users to try out demos. In fact, all users choose a model through the user interface (as described below) and try the demo. The demo simply loads into the server 402 suitable for the demo almost instantly and runs like any other game or application. Whether the demo requires a very high performance server 402 or a low performance server 402, and whatever form of home or office client 415 the user uses, the demo from the user's point of view. Works fine. Software publishers of game or application demos have tight control over what demos are allowed to be tried by users for how long, and, of course, demos have the opportunity to access the full version of the game or application being demonstrated. Can include a user interface element that gives the user.
0288Demos are probably offered for less than a price or for free, so some users try to use the demos (especially game demos that are fun to play repeatedly) repeatedly. Hosting service 210 can use a variety of techniques to limit the use of demos to a given user. The simplest solution is to establish a user ID for each user to limit the number of times a given user ID is allowed to play the demo. However, the user can set a plurality of user IDs when the user ID is free. One technique to address this issue is to limit the number of times a given client 415 is allowed to play the demo. If the client is a standalone device, the device has a serial number and the hosting service 210 can limit the number of times the demo can be accessed by the client with that serial number. When client 415 runs as software on a PC or other device, the hosting service 210 can specify a serial number and store it on the PC, which can be used to limit the use of the demo, but the user can use the PC. A record of the PC network adapter "Media Access Control (MAC)" address (and / or other machine-specific identifiers, such as the hard drive serial number, etc.), if reprogrammable and the serial number can be erased or changed. Another option is provided for the hosting service 210 to retain and limit the use of the demo to it. However, if the MAC address of the network adapter can be changed, this is not a foolproof method. Another solution is to limit the number of times the demo can be played for a given IP address. IP addresses can be redesignated periodically by cable modem and DSL providers, but in practice they are less frequent and IPs block IP addresses for residential DSL or cable modem access. A small number of demo uses can typically be established for a given home if it can be determined to be within (eg, by contacting an ISP). Also, there may be multiple devices behind NAT routers that share the same IP address in the home, but typically in a residential setting, there are only a limited number of such devices. Many demos can be established for a company if the IP address is in a block that serves the company. But in the end, the best way to limit the number of demos on your PC is to combine all the solutions mentioned above. There is no foolproof way for a determined technically skilled user to limit the number of demos that are played repeatedly, but by creating a large number of barriers, troublesome PC users can exploit the demo system. Is also worthless, but rather can create sufficient deterrence so that the demos can be used as intended to try out new games and applications.
0289<u style="single">Benefits of schools, companies and other facilities</u> In particular, significant benefits will be gained for companies, schools and other facilities that use the system shown in Figure 4a. Companies and schools incur significant costs associated with installing, maintaining and upgrading PCs, especially when it comes to PCs for running high-performance applications like Maya. .. As mentioned above, PCs generally utilize only a small portion of the time of the week, and as in the home, the cost of a PC with a given level of performance capability is the office or school environment. Is much higher than the server center environment.
0290For large companies or schools (eg large universities), it is practical for the IT department of such an entity to set up a server center and maintain computers that are remotely accessed over a LAN grade connection. is there. There are many solutions for remote access to computers over a LAN or through a private broadband connection between offices. For example, on a Microsoft Windows terminal server, or through a virtual network computing application like VNC from RealVNC, or Sun Through thin client means from Microsystems, users can gain remote access to a PC or server with a range of qualities in graphic response time and user experience. Moreover, such self-managed server centers are typically dedicated to a single company or school, so heterogeneous applications (eg entertainment and company applications) have the same computational resources at different times of the week. It is not possible to take advantage of the overlapping uses that can be considered when using it. Therefore, many companies and schools lack the scale, resources or expertise to set up a server center on their own with a LAN speed network connection to each user. In fact, most schools and businesses have the same Internet connection as their home (eg, DSL, cable modem).
0291Such tissues still have the need for very high performance computations, either on a stationary basis or on a periodic basis. For example, a small building company has a small number of architects and requires relatively little calculation when doing design work, but periodically requires very high performance 3D calculations (eg). , When creating a new architectural design 3D flythrough for clients). The system shown in Figure 4a is very well suited for such organizations. Organizations do not need more than the same types of network connections (eg, DSL, cable modems) that are typically very cheap and are provided in the home. They either use an inexpensive PC as the client 415, or omit the PC altogether and use an inexpensive dedicated device that easily embodies the control signal logic 413 and the short latency video decompression 412. .. These features are especially appealing to schools that have problems with theft of PCs and damage to dedicated components within PCs.
0292Such a configuration solves a number of problems in such organizations (and many of these effects are also shared by household users who perform common calculations). One of them is that the operating costs (which must ultimately be returned to the user in some way to get a feasible deal) can be very low. This is because (a) compute resources are shared with other applications that have different peak usage times during the week, and (b) organizations access (and incur the cost of) high-performance compute resources only when needed. ), And (c) the organization does not need to provide resources to back up or maintain high performance computing resources.
0293<u style="single">Elimination of piracy</u> Moreover, games, applications, interactive movies, etc. are no longer pirated as they are today. Since the game runs in the service center, the user is not given access to the underlying program code and therefore nothing is pirated. Even if the user copies the source code, the user cannot execute the code on a standard game console or home computer. This opens the market in places around the world like China where standard video games are not available. Also, used games cannot be resold.
0294For game developers, there are few discontinuities in the market today. Hosting Service 210 is a game requirement, as opposed to the current situation where a completely new generation of technology forces users and developers to upgrade and game developers rely on timely delivery of hardware platforms. Can be updated gradually over time as it changes.
0295<u style="single">Streaming interactive video</u> The above explanation is made possible by a new basic concept of short-latency streaming interactive video based on the general Internet (which implies that the audio used here is also included with the video). It describes a wide range of applications. Traditional systems that feed streaming video over the Internet only enable applications that can be embodied with long latency interactivity. For example, playback controls for linear video (eg, pause, rewind, fast forward) work well with long wait times and can be selected from among linear video feeds. And, as mentioned above, the nature of certain video games allows them to be played with long wait times. However, the long latency (or low compression ratio) of traditional solutions for streaming video severely limits potential applications of streaming video or narrows their development to specialized network environments, and Even in such an environment, conventional techniques introduce a substantial burden on the network. The technologies described here open the door to a wide range of possible applications for short-latency streaming interactive video over the Internet, especially those enabled through consumer-grade Internet connections.
0296A client as small as the client 465 in Figure 4c, sufficient to improve the user experience with virtually any amount of computational power, any amount of fast storage, and a very fast network in a powerful server. The device actually enables a new era of computing. Moreover, because bandwidth requirements do not grow as the computing power of the system increases (ie, bandwidth requirements are only tied to display resolution, quality and frame rate), broadband internet connections are unevenly distributed (ie). Typical consumer and enterprise applications when, for example, through widespread short-latency wireless coverage), reliable, and wide enough bandwidth to meet the needs of display devices 422 for all users. On the other hand, the question is whether a thick client (such as a PC or mobile phone running Windows, Linux, OSX, etc.) is required, or a thin client (such as Adobe Flash or Java) may be used.
0297The advent of streaming interactive video has resulted in rethinking assumptions about the structure of computational architecture. One example is the server center embodiment of the hosting service 210 shown in FIG. The video path for the delay buffer and / or group video 1550 is that the multicast streaming bidirectional video output of the app / game server 1521-1525 can be selected in real time via path 1552 or via path 1551. Later, it will be fed back to the app / game server 1521-1525. This enables a wide range of practical applications (eg, as shown in Figures 16, 17 and 20) that are not possible or feasible with traditional servers or local computing architectures. However, as a more general architectural feature, what the feedback loop 1550 gives is iteration at the streaming bidirectional video level. This is because the video can loop back indefinitely when requested by the application. This enables a wide range of application possibilities not previously available.
0298Another important architectural feature is that the video stream is a unidirectional UDP stream. This effectively allows multi-casting of streaming bidirectional video to any degree (in contrast, bidirectional streams of TCP / IP streams are increasingly congested from forward and backward communication to the network as the number of users increases. Will occur). Multicasting is an important capability within the server center. This is because the system enables one-to-many or many-to-many communication in response to the growing needs of Internet users (and in fact the world's population). Again, the example of FIG. 16 showing the use of both streaming bidirectional video iteration and multicasting is only the pinnacle of a very large iceberg of potential.
0299In one embodiment, the various functional modules shown here and related steps are performed by specific hardware components, such as application specific integrated circuits (ASICs), that include fixed wiring logic to perform that step. It is done or performed by a combination of programmed computer components and custom hardware components.
0300In one embodiment, these modules are embodied in a programmable digital signal processor (DSP) of Texas Instruments' TMS320x architecture (eg, TMS320C6000, TMS320C5000, etc.). A variety of different DSPs can be used that fit these basic principles.
0301These embodiments can include the various steps described above. These steps can be performed in machine-executable instructions that allow a general purpose or special purpose processor to perform several steps. Various elements not related to these basic principles, such as computer memory, hard drives, and input devices, have been excluded from the drawings to avoid obscuring the appropriate perspective.
0302The abstract elements disclosed herein may be provided as machine-readable media for storing machine-executable instructions. Machine-readable media include flash memory, optical discs, CD-ROMs, DVD ROMs, RAMs, EPROMs, EEPROMs, magnetic or optical cards, propagation media suitable for storing electronic instructions, or other types of machines. Includes, but is not limited to, readable media. For example, the present invention transfers data signals carried out on a carrier wave or other propagating medium over a communication link (eg, a modem or network connection) from a remote computer (eg, a server) to a requesting computer (eg, a client). Can be downloaded as a computer program.
0303Also, the elements of the gist disclosed herein include machine-readable media containing instructions used to program a computer (eg, a processor or other electronic device) to perform a series of operations. Please understand that it may be offered as a computer program product. Alternatively, the operation may be performed by a combination of hardware and software. Machine-readable media are suitable for storing floppy (registered trademark) discets, optical discs, CD-ROMs, and magnetic / optical discs, ROMs, RAMs, EPROMs, EEPROMs, magnetic or optical cards, and electronic instructions. Including, but not limited to, propagating media or other types of media / machine readable media. For example, the gist element disclosed herein is that a program is transferred from a remote computer or electronic device to a requesting computer by a data signal carried out on a carrier or other propagating medium over a communication link (eg, a modem or network connection). It can be downloaded as a computer program product such as.
0304Furthermore, although the gist disclosed herein has been described in the context of a particular embodiment, numerous changes and amendments may be made within the scope of this disclosure. Therefore, the present invention and the accompanying drawings are merely examples, and the present invention is not limited thereto.
0305100: CPU / GPU 101: RAM 102: Display device 103: hard drive 104: Optical media drive 105: Network connection 106: Game controller 205: Client device 206: Internet 210: Hosting service 211: User's house 220: Software developer 221: Input device 222: Monitor or TV receiver 301: Maximum data rate 302: Maximum data rate actually available 303: Requested data rate 401: Hosting service control system 402: Server 403: SAN 404: Low latency video compression 405: RAID array 406: Control signal 410: Internet 412: Low latency decompression 413: Control signal logic 415: Home or office client 421: Input device 422: Monitor or HDTV 441: Central office, headend, cell tower, etc. 442: WAN interface 443: Firewall / Router / NAT 444: WAN interface 451: Control signal 452: User's house routing 453: User ISP 454: Internet 455: Server Center Routing 456: Frame calculation 457: Video compression 458: Video decompression 462: Power over Ethernet 463: HDMI output 464: Display capability 465: Ethernet vs HDMI Client 466: Glasses with shutters 468: Monitor or SD / HDTV 469: Bluetooth input device 476: Flash 480: Bus 481: Ethernet interface 483: Control CPU 484: bluetooth 486: Video decompressor 487: Video output 488: Audio decompressor 489: Audio output 490: HDMI 497: Ethernet 499: Power
41 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| JP2003152752A | Cites | Japan | Y | Search report | 4-6,29-31 |
| JP2005167515A | Cites | Japan | Y | Search report | 1-50 |
| JP2005244315A | Cites | Japan | Y | Search report | 1-50 |
| JP2005519382A | Cites | Japan | Y | Search report | 1-50 |
| JP2005520265A | Cites | Japan | Y | Search report | 19-21,44-46 |
| JP2006109099A | Cites | Japan | Y | Search report | 1-50 |
| JP2007088539A | Cites | Japan | Y | Search report | 17-18,42-43 |
19 members in 11 offices
Members19
| Document | Office | Kind | |
|---|---|---|---|
| AU2008333833A1 | Australia | A1 | |
| CA2707708A1 | Canada | A1 | |
| WO2009073831A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200941232A | Taiwan Province of China | A | |
| TW200943079A | Taiwan Province of China | A | |
| EP2218016A1 | European Patent Office (EPO) | A1 | |
| KR20100121598A | Republic of Korea | A | |
| CN101918935A | China | A | |
| JP2011507351A | Japan | A | |
| HK1149811A | Hong Kong, China | A | |
| HK1149811A1 | Hong Kong, China | A1 | |
| RU2010127307A | Russian Federation | A | |
| EP2218016A4 | European Patent Office (EPO) | A4 | |
| NZ585906A | New Zealand | A | |
| RU2493585C2 | Russian Federation | C2 | |
| JP2013211902A | Japan | A | |
| CN101918935B | China | B | |
| JP2017076984AThis record | Japan | A | |
| JP6442461B2 | Japan | B2 |
16 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 2017076984
- Application
- 213073
Titles2
- Japanese
- 通信チャンネルにわたるパケットロスの影響を減少するためのビデオ圧縮システム及び方法
- English
- Video compression systems and methods to reduce the effects of packet loss across communication channels
Classification
- CPC, 27
- A63F13/12
- A63F13/358
- A63F2300/538
- A63F2300/577
- H04N21/2343
- H04N21/2383
- H04N21/2402
- H04N21/2662
- H04N21/4781
- H04N21/6587
- H04N19/30
- H04N19/61
- H04N19/114
- H04N19/132
- H04N19/14
- H04N19/137
- H04N19/146
- H04N19/166
- H04N19/17
- H04N19/436
- H04N19/587
- H04N19/59
- H04N19/188
- H04N19/88
- A63F13/30
- A63F13/335
- A63F13/77
- IPC, 3
- H04N21 238
- H04N21 23
- H04N19 89