Video generation based on text
Abstract
Solution.A method of generating a video string of a person based on a character string is disclosed. The method is a step of generating a human video sequence by the processing device in order to model a visual and audible human emotional expression based on the reception of a character string. A method comprising the step of using a voice model of the person's voice to generate a voice portion of the person. The emotional expression in the visual portion of the video sequence is modeled on the prior knowledge of the person. For example, prior knowledge includes photographs and videos of the person taken in real life. [Selection diagram] Fig. 10

Term
5.6 yearsto projected expiry
Projected expiry 4 May 2032, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
24 claims: 8 independent, 16 dependent
- 1処理装置に文字列を入力するステップと、 視覚的かつ可聴的な、人の感情表現をモデル化するために、前記処理装置により、前記文字列に基づいて前記人の映像列を生成するステップであって、前記映像列の音声部分を生成するために、前記人の音声の音声モデルを用いることを有するステップとを含む方法。
- 2前記処理装置は、携帯装置であり、前記文字列は、ショートメッセージサービス(SMS)を介して第二携帯装置から入力され、 人の映像列を生成する前記ステップは、前記携帯装置と前記第二携帯装置に記録された共有情報に基づいて人の映像列を前記携帯装置により生成することを含む請求項1に記載の方法。
- 3前記文字列は、少なくとも一つの言葉を含む言葉群を有し、前記映像列は、前記人が映像列の中で言葉を発声して見えるように生成される請求項1に記載の方法。
- 4前記文字列は、発声を表現する文字を有し、前記映像列は、前記人が映像列の中で言葉を発声して見えるように生成される請求項1に記載の方法。
- 5前記文字列は、言葉と該言葉の指標とを有しており、前記指標は、前記映像列の中で前記人が前記言葉を発声して見えるとき、前記映像列の中で前記人の感情表現を同時に表示し、前記指標は、規定の指標群であり、前記規定の指標群の各指標は、異なる感情表現に関連する請求項1に記載の方法。
- 6映像列を生成する前記ステップは、前記文字列と前記人の事前知識に基づいて、視覚的かつ可聴的な、前記人の感情表現をモデル化するために、前記処理装置により人の映像列を生成することを含む請求項1に記載の方法。
- 7映像列を生成する前記ステップは、前記文字列の言葉を前記人の顔の特徴にマッピングするステップと、前記人の顔の特徴を背景上にレンダリングするステップとを含む請求項1に記載の方法。
- 8前記言葉は、前記言葉のための一又は複数の指標に基づいて前記顔の特徴にマッピングされ、前記指標は、前記映像列の中で前記人が前記言葉を発声して見えるとき、前記映像列の中で前記人の感情表現を同時に表示する請求項7に記載の方法。
- 9前記顔の特徴は、特定の前記人に適用される特定の顔の特徴を含む請求項7に記載の方法。
- 10映像列を生成する前記ステップは、前記人の顔の特徴に適合する前記人の体のジェスチャーを生成することを含む請求項7に記載の方法。
- 11映像列を生成する前記ステップは、前記人の音声に基づく音声モデルを用いて、前記文字列内の言葉に基づいて前記人の音声を表現する音声列を生成することを含む請求項1に記載の方法。
- 12文字列の受信は、リアルタイムで文字列を受信することを含み、 映像列を生成する前記ステップは、視覚的かつ可聴的な、前記人の感情表現をモデル化するために、前記文字列に基づいて人の映像列をリアルタイムで生成するステップを含み、該ステップは、前記映像列の音声部分を生成するために前記人の音声の音声モデルを用いることを含む請求項1に記載の方法。
- 13処理装置に文字列を入力するステップと、 視覚的な、人の感情表現をモデル化するために、前記処理装置により、前記文字列に基づいて前記人の映像列を生成するステップであって、前記映像列の各フレームの顔部が前記人の複数の推測画像の結合により表現されるステップと、 可聴的な、人の感情表現をモデル化するために、前記人の音声の音声モデルを用いて、前記文字列に基づいて前記人の音声列を前記処理装置により生成するステップと、 前記処理装置を用いて、前記映像列と音声列とを結合することにより前記人の映像列を生成するステップであって、前記映像列と音声列が前記文字列に基づいて同期されるステップとを含む方法。
- 14前記映像列の各フレームの顔部は、前記人の複数の推測画像の線形結合により表現され、前記人の複数の推測画像における各推測画像は、前記人の平均画像からの偏差に対応する請求項13に記載の方法。
- 15前記文字列に基づいて前記人の映像列を生成する前記ステップは、前記映像列の各フレームを2以上の領域に分割することを含み、少なくとも一つの前記領域は、前記人の推測画像の結合により表現されている請求項13に記載の方法。
- 16前記人の音声の前記音声モデルは、前記人の音声サンプルから作られる複数の音声の特徴を含み、前記複数の音声の特徴の各音声の特徴は、文字に対応する請求項13に記載の方法。
- 17前記複数の音声の特徴における各音声の特徴は、言葉、又は、音素、発声に対応する請求項16に記載の方法。
- 18前記人の音声の音声モデルは、前記人の音声サンプルから作られる複数の音声の特徴と、前記文字列に従う第二の人の音声と、前記人の音声の波形と第二の人の音声の波形との一致とを含み、前記人の音声特徴は、前記人の音声の波形と第二の人の音声の波形との前記一致に基づいて、前記第二の人の音声にマッピングされる請求項13に記載の方法。
- 19前記人の音声の音声モデルは、前記人の音声サンプルから作られる複数の音声の特徴と、前記文字列に従って、文字を音声に変換するモデルにより生成される音声と、前記人の音声の波形と文字を音声に変換するモデルの音声の波形との一致とを含み、前記人の音声特徴は、前記人の音声の波形と文字を音声に変換するモデルの音声の波形との前記一致に基づいて、前記モデルの音声にマッピングされる請求項13に記載の方法。
- 20文字列を生成するステップであって、前記文字列は、人の感情の範囲を視覚的かつ可聴的に表現するために、前記人の音声に基づく音声モデルを用いて生成される映像列の中で人が発声する一又は複数の言葉を表現するように構成されるステップと、 前記文字列内の言葉に関連する指標を特定するステップであって、前記指標は、規定の指標群の一つであり、各指標が前記人の異なる感情表現を示すように構成されるステップと、 前記指標を前記文字列に組み込むステップと、 前記映像列を生成するように構成される装置に前記文字列を送信するステップとを含む方法。
- 21指標を特定する前記ステップは、前記文字列内の言葉に関連する項目群の一覧から一項目を選択するステップを含み、前記一覧の各項目は、前記人の感情表現を示す指標である請求項20に記載の方法。
- 22指標を特定する前記ステップは、自動音声認識(ASR)装置を用いて、前記文字列内の言葉を話す話者の音声列に基づいて、前記文字列内の言葉に関連する指標を特定するステップを含む請求項20に記載の方法。
- 23非個人の複数の項目の情報を処理装置に記憶するステップと、 前記処理装置の前記非個人の複数の項目の事前情報に基づいて、前記非個人の複数の項目のための映像列を生成するステップであって、前記非個人の各項目が独立して管理可能に構成されるステップとを含む方法。
- 24前記非個人の複数の項目は、前記映像列の中で他の要素に関連して制約される請求項23に記載の方法。
Independent claims24
121 paragraphs, as filed
Reference of related application
This application enjoys the benefits of US provisional application 61 / 483,571 filed May 6, 2011, which is incorporated herein by reference.
At least one embodiment of the present application relates to video generation, more specifically, a visual and audible person based on a character string (character sequence) such as a user-to-user short text message. It relates to a method and a system for generating a video (video) in order to realistically model (simulate) the emotional expression of.
The digital video sequence (video sequence) contains a large amount of data. Large data transmission bandwidths are required to efficiently transmit video using current technology. However, the data transmission bandwidth for wireless media is limited and it is also an expensive resource.
For example, the Short Message Service (SMS), sometimes referred to as mobile email, is one of the most well-known interpersonal message technologies used today. The SMS feature is widely available on almost all modern mobile phones. However, SMS has a very limited capacity for transmitting information, and each SMS message has a fixed length of 140 bytes or 160 characters, and is therefore not suitable for transmitting video data. Multimedia Message Service (MMS) is a method by which it is possible to send a message containing multimedia content. However, MMS messages do not take advantage of the existing SMS infrastructure and are more expensive than SMS messages. Sending video messages on very low (small) bandwidth channels, such as SMS channels, is difficult, if not impossible.
<p> The techniques of the present application include methods and devices for generating a video sequence of a person based on the character string in order to realistically model the person speaking according to the character string. The technique involves realistically modeling a person's emotional expression, which is visible and audible based on a character string. The technique can create the appearance of transmitting a person's video on a low bandwidth channel (eg, an SMS channel) without actually transmitting the video data.</p><p> Moreover, the technique is a cost-effective way to produce realistic footage that is not a video of the person speaking (or singing or making any vocalizations). The technique provides an alternative method of recording a person's footage when the schedule is uncoordinated or reluctant, or inconvenient for various reasons such as death or emergency situations. In addition to the specific prior information of the person recorded on the processing device, the technique only requires a string to be able to generate the video, and the string is transmitted if it needed to be transmitted in the past. Does not need. The string provides a mechanism for controlling and adjusting the words and emotional expressions so that the person appears to be uttered in the video string by adjusting the content of the string.</p><p> In one embodiment, the processing device receives a character string. Based on the received string, the processing device includes using an audio model of human audio to generate an audio portion of the video sequence to model a visible and audible human emotional expression. Generate a video string of people. The emotional expression in the audio part of the video sequence is modeled on the basis of prior knowledge of the person. For example, prior knowledge includes photographs of people taken in real life and / or images of people.</p><p> In certain embodiments, the string comprises a word group (a group of words) containing at least one word. The video sequence is generated so that a person can speak and appear in the video sequence. The string further includes emotional indicators related to one or more words. Each index simultaneously displays a person's emotional expression in the video sequence when a person appears to utter a word for the index in the video sequence. In certain embodiments, the processor maps words in a string to features of a person's face, based on prior knowledge of the person. The processing device makes a feature of a human face a background image.</p><p> Other features of the technology of the present application will be clarified from the drawings and the detailed description below.</p><p> The above and other objects, configurations and properties of the present invention will be more apparent to those skilled in the art through claims, drawings and the following detailed description.</p>
<figref num="1A">An example of a processing device that receives a character string from a transmitting device is shown.</figref><figref num="1B">An example of a processing device that receives a character string from a transmitting device via an intermediary device is shown.</figref><figref num="2">An example of a processing device that receives a character string from the input unit of the processing device is shown.</figref><figref num="3">A high-level block diagram showing a configuration example of the processing device is shown.</figref><figref num="4">An example of the configuration of a system that converts characters into video (TTV) is shown.</figref><figref num="5">An example of constructing a human video model is shown.</figref><figref num="6">An example of the process of dividing the target person's face into two areas is shown.</figref><figref num="7">An example of dictionary registration (entry) is shown.</figref><figref num="8">An example of the process of creating a model that converts characters into speech (TTS) is shown.</figref><figref num="9">An example of the TTS speech synthesis process is shown.</figref><figref num="10">An example of the process of synthesizing a video string and an audio string is shown.</figref><figref num="11">An example of the process of embedding a composite video in the background while minimizing the boundary matching error is shown.</figref><figref num="12">An example of the process of embedding a composite video in the background based on a two-region model is shown.</figref>
References such as embodiments herein mean that the particular features, configurations, or properties described are included in at least one embodiment of the invention. All such references herein do not necessarily refer to the same embodiment.
Some related references are "Neighborhood Fits of Lines and Planes to a System of Spatial Points" (Book Philosophical Magazine 2 (6), pp. 559-572, 1901, author K. Pearson), "Appearance of Computer Screens". Statistical Model (Book Technical Report, University of Manchester, pp. 125, 2004, Author TF Cootes, CJ Taylor), "Dynamic Appearance Model" (Book Proc. European Conf. Computer Vision, 2, pp. 484-489, 1998) Year, author T. Cootes, G. Edwards, C. Taylor), "Proceedings of the 26th annual conference on Computer graphics and interactive techniques, pp. 187-194, publisher ACM Press / Addison-" Wesley Publishing Co., 1999, Book V. Blanz, T. Vetter) "Face emotions" (Author Stanford Computer Science Technical Report, CSTR 2003-02, Author Erica Chang, Chris Bregler) "Real-time human mouth movements driven by human voice" (Author IEEE Workshop on Multimedia Signal Processing, 1998, author FJ Huang, T. Chen), "Extracting Voice Features for Lip Reading" (Author IEEE Transactions on Pattern Analysis and Machine Intelligence, 2002 24 (2), pp. 198-213, Author Matthews, I. , Cootes, T., Bangham, A., Cox, S., Harvery, R.), "Lip Reading Using Unique Rows" (Book Proc. Int. Workshop Automatic Face Gesture Recognition, 1995, pp. 30-34, Author N. Li, S. Dettmer, M. Shah), "Movement of Video Speech with Audio" (Book Proceedings of SIGGRAPH 97, pp. 353-360, August 1997, author Christoph Bregler, Michele Covell, Malcolm Slaney). All of them are incorporated herein by reference.
FIG. 1A shows a processing device and an environment in which the technology introduced in the present application is executed. In FIG. 1A, the processing device 100 is connected to the transmitting device 110 via the interconnect 120. The interconnection 120 is a mobile phone network or a global area network such as an SMS channel, a television channel, a local area network (LAN), a wide area network (WAN), an urban scale network (MAN), or the Internet. , Fiber channel structure, a combination of such interconnects. The transmitting device 110 can transmit the character string 140 to the processing device via the interconnection 120. The processing device 100 receives the character string 140 and generates the video string 150 based on the character string 140. Either the transmission device 110 or the processing device 100 is, for example, a mobile phone, a conventional personal computer (PC), a server-class computer, a workstation, a portable computer / communication device, a game machine, a television, or the like.
The processing device 100 has a storage device 160 that stores the generated video sequence 150. The storage device 160 is, for example, a conventional dynamic random access memory (DRAM), a conventional magnetic disk, an optical disk, a tape device, a non-volatile semiconductor memory such as a flash memory, a combination of the above devices, and the like.
One of the processing device 100 and the transmitting device 110 has operating systems 101 and 111 that manage (control) the operations of the processing device 100 and the transmitting device 110. In certain embodiments, the operating systems 101,111 are run in software. In other embodiments, one or more of the operating systems 101,111 are run on genuine hardware, for example as a specially designed dedicated circuit or as a partial dedicated circuit of software.
A character string such as the character string 140 in FIG. 1A has an index (also called a tag, or an emotional index or an emotional tag). Each index simultaneously indicates a person's emotional expression in the video sequence when the person appears to utter a word in the video sequence. The indicators have different configurations and are selected by different methods. In one embodiment, the index is selected as one item from a list of a plurality of items related to a word in a character string, and each item in the list is an index indicating a person's emotional expression. In other embodiments, the indicator is identified by inserting a markup language string used for the words in the string. The markup language character string is composed of a group of default markup language character strings, and each markup language character string in the group is an index indicating a person's emotional expression. In yet another embodiment, the index is identified by the speech sequence of the speaker who speaks the words in the string using an automatic speech recognition (ASR) device.
FIG. 1B shows an example of a processing device that receives a character string from a transmitting device via an intermediate (mediating) device. In FIG. 1B, the processing apparatus 100 is connected to the intermediate apparatus 180 via the interconnect 192. The transmitter 110 is connected to the intermediate device 180 via the interconnect 191. Either the interconnect 191 and the interconnect 192 may be, for example, a mobile phone network or an SMS channel, a television channel, a local area network (LAN), a wide area network (WAN), an urban network (MAN), or the Internet. Global area networks such as fiber channel structures, combinations of such interconnects. In some embodiments, the interconnect 191 and the interconnect 192 are in one network, such as the Internet. The transmitting device 110 can transmit the character string 140 to the intermediate device 180 via the interconnection 191. The intermediate device 180 further transmits the character string 140 to the processing device 100 via the interconnect 192. The processing device 100 receives the character string 140 and generates the video string 150 based on the character string 140. The intermediate device is, for example, a mobile phone, a conventional personal computer (PC), a server-class computer, a workstation, a portable computer / communication device, a game machine, a television, or the like.
In some embodiments, the intermediate server 180 receives the string and processes the string 140 in the dataset. The data set is transmitted to the processing device 100 instead of the character string 140.
FIG. 2 shows an example of a processing device that receives a character string from the input unit of the processing device. The processing device 200 has an input unit 210 capable of receiving the character string 240 from the person 290. The processing device 200 is, for example, a mobile phone, a conventional personal computer (PC), a server-class computer, a workstation, a portable computer / communication device, a game machine, a television, or the like. The input unit 210 is, for example, a keyboard, a mouse, an image / video camera, a microphone, a game machine controller, a remote controller, a sensor, a scanner, a music device, or a combination of such devices.
The processing device further comprises a processor 205 that generates a human video string 250 based on a character string 240 and a human prior knowledge 270. The video sequence 250 models a visible and audible expression of a person's emotions, and the person appears to utter a specific word in the character string 240 in the video sequence 250. The generated video sequence 250 is stored in the storage device 260 in the processing device 200. The storage device 260 is, for example, a conventional dynamic random access memory (DRAM), a conventional magnetic disk, an optical disk, a tape device, a non-volatile semiconductor memory such as a flash memory, a combination of the above devices, and the like. The character string 240 and / or the person's prior knowledge 270 is stored in the storage device 260 or in another storage device away from the storage device 260.
The processing device 200 has an operation system 201 that manages the operation of the processing device 200. In certain embodiments, the operating system 201 is run in software. In another embodiment, the operation system 201, a dedicated circuit designed for example specially, or, Sofutou such as a partial dedicated circuit E A, is executed in the genuine hardware.
FIG. 3 shows a block diagram of a processing device used to perform the above techniques. In certain embodiments, at least some of the members shown in FIG. 3 are distributed between two or more computer platforms or computer boxes that are connected apart from each other. The processing device is a conventional server class computer, or a PC, a portable communication device (for example, a smartphone), a tablet computer, a game machine, or other well-known or conventional processing / communication device.
The processing device 301 of FIG. 3 has one or more processors 310. The processor 310 is, for example, a central processing unit (CPU), a memory 320, an Ethernet® adapter, and / or at least one communication device such as a wireless communication system (eg, cellular, WiFi, Bluetooth, etc.). The 340 and one or more I / O devices 370,380 are connected to each other via an interconnect 390.
The processor 310 manages the operation of the processing device 301. The processor 310 comprises one or more programmable general purpose or special purpose microprocessors or a combination of microcontrollers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), and the above-mentioned equipment. Or have. The interconnect 390 has one or more buses, direct connections, and / or other types of physical connections, and various bridges, controllers, and / or adapters that are well known to those of skill in the art. Have. The interconnect 390 further has a system bus. The system bus is configured to be connected to one or more expansion buses through one or more adapters, a peripheral component interconnect (PCI) bus, or a hypertransport or industry standard architecture (ISA) bus. It has a small computer system interface (SCSI) bus, a universal serial bus (USB), and an IEEE 1394 standard bus (also called a firewire).
The memory 320 comprises or has one or more storage devices of one or more types such as read-only memory (ROM), random access memory (RAM), flash memory, disk drive, and the like. The network adapter 340 is a device suitable for allowing the processing device 301 to communicate with data using a remote device over a communication line, for example, a conventional telephone modem, a wireless modem, a digital subscriber line (DSL) modem, or a cable modem. , Wireless transmitters / receivers, satellite transmitters / receivers, Ethernet adapters, etc. The I / O devices 370,380 include one or more devices such as, for example, a mouse, a trackball, a joystick, a touchpad, a keyboard, a microphone having a voice recognition interface, a voice speaker, a pointing device such as a display device, and the like. However, the I / O device does not have to be provided in a system that operates as a dedicated server and, like the servers of at least some embodiments, does not have a direct user interface. Other forms of the indicated member groups are performed in a manner consistent with the present invention.
The software and / or firmware 330 that programs the processor 310 that performs the above operations is recorded in memory 320. In certain embodiments, such software or firmware is provided to processing device 301 by downloading from a remote system through processing device 301 (eg, via a network adapter 320).
FIG. 4 shows a configuration example of a character-to-video (TTV) system. The system 400 for converting characters into images is executed on one processing device, or on a group of processing devices, and / or on a server. The system 400 has a video database 410 containing a human video model. The person is referred to as a "target person", or "target individual", or simply a "target", and the face thereof is referred to as a "target face". The video model includes, for example, prior information such as an image or video of the target person. After receiving the character 430, the system 400 creates a dictionary. The dictionary maps characters 430 to the movement of the target person's face based on the target person's video model. The video sequence of the target person is created based on the dictionary. In some embodiments, the information from the referenced individual is used to create a video sequence, as detailed in the following paragraphs. The background scene is created in the video sequence, and the target face is overlaid on the top layer of the background scene.
System 400 has a voice database 420 containing the target person's voice model. The voice model has prior information about the target person and the referenced individual. The details of this approach are explained in the following paragraphs. The video sequence and the audio sequence are combined and fused (combined) into the target person's video sequence (450). In some embodiments, the video sequence is output on a display or to a remote device (460).
The video model of the target person will be explained.
FIG. 5 shows an example of constructing a video model of the target person, especially the target person's face. The data of the video model is recorded in a processing device that executes video generation based on characters. In one embodiment, the video model is created by taking one or more sample footage of the person while the person is speaking a language. The number of words spoken by the person needs to be so large that a rich facial expression of the person's face, lips, and mouth is captured in the image. At the stage of video modeling, the spoken words do not have to be associated with the words supplied at a later stage of video generation. As a general assumption, no prior knowledge of the characters to be supplied or input for video generation is required.
All that is needed to build a video model is sufficient information about mouth and face movements. In some embodiments, sample footage in different languages is used to build the footage model. In one embodiment, for example, the required training data includes 5 minutes of video. When constructing a video model, a representative frame from the video that captures typical facial movements is selected as a feature point to build the model (510). Feature points are classified manually or automatically. These feature points indicate important or typical facial features of a person (for example, when the upper lip and lower lip meet, or when the upper eyelid and lower eyelid meet), and the above-mentioned important points. Includes midpoints between facial features.
For each selected frame, the N points of the frame are selected as mesh points to characterize the person's face. Therefore, each frame is defined in 2N-dimensional Euclidean space coordinates (each point is indicated by x and y coordinates). These points represent the facial shapes of individuals with different emotions, so they are not randomly scattered in higher dimensional space. In one embodiment, a dimension reduction method such as Principal Component Analysis (PCA) is applied to these points. For example, a linear dimension reduction method such as PCA is applied. An ellipsoid is defined for the mean image of the face, and the principal axis is defined as the eigenvector of the data autocorrelation matrix. These principal axes are limited according to the magnitude of the eigenvectors of the autocorrelation matrix. The eigenvector with the maximum eigenvalue indicates the direction with the greatest volatility among the N points. Each eigenvector of the small eigenvalues indicates a small volatility and a less important direction. In one embodiment, the K largest eigenvectors are sufficient to realistically represent every possible movement of the face. Therefore, the movement of each face is expressed as a group of K, and the number is called a multiplier. Each multiplier indicates the degree of expansion from the mean image along the direction of the corresponding eigenvectors among the K most important eigenvectors. These eigenvectors are called shape eigenvectors. The shape eigenvectors form the human shape model 520. In some embodiments, the K number is applied and adjusted based on the type of video being processed.
To represent the pixel color of the face, the mesh points of the average image are used to make a triangulation (make a triangulation) of the average image. In the triangulation (triangulation) process, the face image is divided into multiple triangular regions, and each triangular region is defined by three mesh points. Corresponding triangle divisions are made based on the displaced mesh points (relative to the mesh points in the average image) for other facial movements derived from the average image. In one embodiment, a mesh point triangulation process is performed for each of the N frames. These triangulations are used to create a linear mapping from each triangulation region within the classified frame to the corresponding triangulation region of the mean image. The pixel values of the N classified images are moved to the image defined inside the boundaries of the average shape.
PCA is performed on these images defined in the inner region of the average image. A large number of images are retained after PCA to represent facial textures. These retained images are called texture eigenvectors. The texture eigenvectors form a human texture model. Similar to the multiplier for the shape eigenvectors, the multiplier for the texture eigenvectors is used to represent the pixel color of the face (in other words, the texture). The group of multipliers for the shape eigenvectors and the texture eigenvectors is used to reproduce the target person's face 540. In some embodiments, for example, the total number of eigenvectors (or corresponding multipliers) is approximately 40-50. In the rendering device (processing device), each frame of face movement is reproduced by a linear combination of a shape eigenvector and a texture eigenvector using a multiplier as a linear coefficient. In one embodiment, the eigenvectors are recorded in the renderer.
The shape model division will be described.
In some embodiments, the target person's face is divided into multiple areas. For example, FIG. 6 shows an example of a process in which a target person's face is divided into two regions, an upper region and a lower region. Separate sets of shape eigenvectors 610,620 are used to model the lower and upper regions. A separate set of multipliers 614,624 used for the shape eigenvectors 610,620 is used to model the lower and upper regions based on the lower and upper face shape models 612,620. The synthetic lower region 616 represented by the multiplier 614 is combined with the synthetic upper region 626 represented by the multiplier 624 to generate the entire synthesized face of the target person. The lower region is of greater interest than the upper region for video generation of the speaker. Therefore, the lower region is represented by more multipliers and eigenvector pairs than the upper region.
The video model of the referenced individual will be described.
In some embodiments, a video model of the reference individual is created for the target person using a technique similar to that described above. In one embodiment, these reference individual video models are larger than the target person because these reference individual models are used to create a dictionary that maps text content to facial movements. Made using a set. The story (speech) content is large enough to reproduce almost any possible movement that results from a typical story in different emotions.
A dictionary that maps characters to movements will be described.
One or more dictionaries are created to map textual content to facial movements based on the reference individual's video model. In one embodiment, the possible textual content is broken down into words, phonemes, and vocalizations. Vocalization is a sound that cannot be expressed in words. Each word, phoneme, and utterance has at least one entry in the dictionary. In some embodiments, the word, or phoneme, utterance has multiple registrations in the dictionary. For example, a word has multiple registrations in a dictionary corresponding to different emotions.
In one embodiment, during the creation of a video model for a reference individual, the reference model speaks the most common, many words, in a language for humans to see in the generated video. In other words, words with different emotions are reproduced based on constituent phonemes using information from footage recorded for the referenced individual.
Each registration in the dictionary is a mapping between words or phonemes, utterances and time series shape multipliers (time series are called frame series (continuous) or disclosure columns). For example, if it takes a period of T frames for the reference individual to make a facial movement that says the word "rain", then the time series of shape multipliers is f (k, t), k = 1 ~ K, It is indicated by t = 1 ~ T. In each frame t, the movement of the reference individual's face is represented by K numbers of the shape multiplier f (k, t). Thus, a group of total K × T multipliers represents a series of facial movements of the reference individual corresponding to the word rain. That is, the registration in the dictionary is as follows. "Rain": f (k, t), k = 1 ~ K, t = 1 ~ T
In one embodiment, the registrations in the dictionary are automatically accumulated using an automatic speech recognition (ASR) device. The ASR device can recognize both words and phonemes. In some embodiments, the ASR device can further recognize utterances that are not words or phonemes. If the word "rain" is spoken with different emotions, the dictionary can contain multiple registrations for the word "rain" with different emotions. For example, in one registration, it is as follows. "Rain (surprise emotion)": f<sub>1</sub>(k, t), k = 1 ~ K, t = 1 ~ T
In some embodiments, some words are composed of phonemes. The time series of multipliers for phonemes depends not only on the phonemes, but also on adjacent phonemes (or silence before and after the phonemes) uttered before and after. Therefore, the dictionary can include a plurality of registrations of phonemes. When registration is used to generate a video sequence, the choice of phoneme registration depends on the adjacent phonemes (or silence before and after the phoneme) within the supplied string.
FIG. 7 is an example of dictionary registration 710. Dictionary registration 710 maps the word 712 to a time series of multipliers 714. Each frame 720 of facial movement is represented by a group of multipliers from a time series of multipliers 714. ASR device 730 is used for the accumulation of registration 710.
The optimization of the dictionary by vector quantization will be described.
In some embodiments, when constructing a dictionary, some words, phonemes contained in the dictionary are uttered many times by the reference individual. This makes it possible to generate a dictionary containing many registrations for one word or phoneme, and each registration maps words or phonemes to time series with different multipliers. For example, as mentioned above, in the case of phonemes, the selection of dictionary registration for phonemes is based on history (ie, the shape of the mouth for a particular phoneme is due to the adjacent phonemes uttered before and after it. It will be decided.). This choice is too many, i.e. the dictionary provides many choices of registration.
To improve predictability and search efficiently, dictionaries are optimized in the following ways: The video sequence for audio can be thought of as a large number of points in the space of multipliers. For example, if a 30-minute video, which is 30 frames per second, is used, it will have a group of multiplier values of 30 × 1800 × K = 54000 × K. Here, K is the number of shape eigenvectors used in the video model. Some of these points represent mouth positions that are very close to each other.
Vector quantization (VQ) is performed on a population of 54,000 points in K-dimensional space. In VQ, 54000 points are approximated by M centers of the population (VQ points, or VQ centers, VQ indicators). Here, each point is replaced by the closest center of the population. The larger the number of centers, the better the reproducibility because the VQ points are 54000 points. This is because the movement of the face is very restricted to the point group. There is a correlation between the multipliers. Therefore, positive VQ reproduction is possible. In one embodiment, the number of VQ centers is determined so that the maximum error is tolerated. Here, the maximum error is visually confirmed to provide acceptable behavior of the video.
Thus, after vector quantization, the dictionary registration includes the time series at the center of the VQ instead of the time series of multipliers for the corresponding eigenvectors. This allows for a more compact registration of dictionaries to express words and phonemes. Multiple registrations of words and phonemes in the first dictionary break less registrations, including the timeline at the center of the VQ (the same utterances are made different times by the reference speaker, so they are in the same timeline at the center of the VQ. Can be mapped.). Moreover, this destruction makes it more manageable to choose a time series of phonemes based on adjacent phonemes.
The voice model will be described.
The purpose of the speech model is to create an arbitrary sentence or a group of sentences based on a given character. Several techniques for creating speech-based speech models are disclosed in the following paragraphs.
A voice model (TTS) that converts characters into voice will be described.
Figure 8 shows an example of the process of creating a TTS speech model. The TTS speech model collects speech samples 810 based on the text file 840 for the target person to create the speech model. The voice data in the voice sample is used to create a group of voice features of the target person. In one embodiment, the audio feature has an excited state element 820 and spectral information 830. These audio features and the corresponding extracted text content 850 are inputs used to create audio models 860,870 and improve the accuracy of the model. Once the speech model is created, a new string is supplied to generate the speech. This speech model is a probabilistic model, i.e., given a new string, a group of speech features from the speech model are combined with the speech string to plausibly represent the new string.
For example, FIG. 9 shows an example of the TTS speech synthesis process. The character 920 is input to the establishment model 910. In the 930, the parameter string representing the voice feature is selected by the model for representing the character string 910. Speech parameters that represent speech features are transformed by the model that produces the speech waveform, thus synthesizing speech sequences (940,950).
The output of the TTS system includes not only the speech sequence of words and phonemes, but also its time marker (also referred to as a time stamp). For example, consider the word "rain" to be part of a character that is converted into a speech sequence. The audio model not only generates the word "rain", but also generates the first and last time stamps for the "rain" audio sequence for the first time of the generated audio sequence. This time stamp is used for audio-video synchronization disclosed in the following paragraphs.
The direct TTS model, which synthesizes characters into speech, creates speech that is directly related to the speech data used to generate the model. The advantage of this technique is that once the model is created, the generation of spoken speech requires only spoken characters.
The voice conversion voice model will be described.
Other techniques for creating speech models are based on building a match (correspondence) between the voices of two speakers. One speaker is the reference speaker and the other speaker is the target speaker. In this technique, voice data based on the same character is collected from the target speaker and the reference speaker. A match (correspondence) is established between the acoustic waveforms of the reference speaker and the target speaker. The match is used to generate the voice of the new word of the target speaker based on the voice of the new word spoken by the reference speaker.
The match between the reference voice and the target voice is constructed by the following method. Voice samples from target and reference speakers who speak the same language are collected. In one embodiment, the audio sample has a length of several minutes. By analyzing this waveform, the utterances of the reference voice and the target voice are adjusted (aligned) so that the utterances of the reference voice and the target voice match. The voice characteristics of the reference voice and the target voice (such as the mel frequency cepstrum coefficient) are extracted. This binding distribution is modeled by GMM (Gaussian mixed model). The first estimation of GMM parameters is made by vector quantization of feature clusters in the coupling histogram. GMM is trained by the EM (Expected Value Maximization) algorithm.
Using this technique, the features of the reference voice are mapped to the features corresponding to the goal. From these corresponding features, the acoustic waveform is generated as a voice sequence of the target person. In some embodiments, feature alignment becomes noise at the first step of the process. The generated target voice (as opposed to the first target voice) is assigned to the algorithm as input for repeated execution until convergence.
This voice conversion model has several advantages. The first advantage is that the emotional state of the voice is transmitted from the reference to the target. The second advantage is that when the referenced video is made for audio, it facilitates the high quality video rendering (video representation) of the target. Thus, for example, a speech conversion model is useful for entertainment purposes when a particularly rigorous emotional effect of the goal is required.
Describes PCA-based voice conversion.
Basic GMM-based voice conversion improves effectiveness and speed by utilizing PCA (Principal Component Analysis). In this case, GMM voice conversion training is performed with one reference voice and multiple target voices. Multiple GMMs trained with different target speeches are put into the PCA process to break down speech volatility.
If the generated target audio sample is very large, adding a new target audio does not require the collection of multiple audio samples and training of a new GMM. Instead, only short-time audio samples of the new goal are obtained, and their GMM parameters are determined by the decomposition of the PCA eigenvectors based on the previously trained GMM. With a sufficient original source (first resource) training set to multiple GMMs for different goals, the quality of the generated speech is because the PCA eliminates noise volatility within a single GMM process. , Will be improved.
To summarize the technique, the reference data is transformed into multiple trained target GMMs. The PCT model is generated for multiple trained target GMMs. For the new goal person, the PCA decomposition is performed to synthesize a speech sequence for the new goal person, where only limited training data is the new goal. Required by people.
Describes TTS-based PCA voice conversion.
The above reference voice does not have to be a natural human voice. It may be a voice generated in a high quality TTS. This TTS-generated voice does not have to be the voice of a particular individual. The exact same process as described above is performed, except that the high quality synthetic TTS instead of the reference individual voice is the reference voice.
The advantage of using fixed synthetic TTS resources is that there is no need to go back to the human voice to generate the voice resource of the new set of words in order to generate the voice of the new goal. Therefore, only a character string is required as an input for video generation.
Audio-video synchronization will be described.
The synthesis of the generated video and audio sequences is performed in different ways, depending on whether the direct TTS or voice conversion method was used to create the synthetic audio. This method requires building a relationship between the reference individual and the video model of the target person. The relationship is generated by the alignment of the reference and the shape eigenvectors of the target. The alignment is represented by a transformation matrix. The transformation matrix requires only one calculation and is recorded in the rendering device. The size of this conversion matrix is small. It is assumed that the target person is represented by 20 shape eigenvectors and 23 texture eigenvectors, and the dictionary is accumulated based on a reference individual represented by 18 shape eigenvectors and 25 texture eigenvectors. .. That is, the transformation matrix is a 20x16 and 23x25 matrix for each of the shape eigenvectors and the texture eigenvectors. Only this matrix is recorded in the renderer for conversion purposes. Accumulation of a database of reference individuals used to create the dictionary is unnecessary on the rendering device.
The synthesis of the voice conversion voice model will be described.
When the target audio is generated by the voice conversion method, the process of synthesizing with the video is carried out as follows. Audio is created on the basis of reference individuals for whom we have video and audio models. The multiplier of the shape eigenvector and the texture eigenvector of the reference individual is calculated. This multiplier is converted into a multiplier of the shape eigenvector and the texture eigenvector in order to generate the video sequence of the target person.
The generated video sequence of the target person has face and lip movements synthesized with the video sequence of the target person, and any emotions displayed in the video sequence. Thus, the target person's video sequence is provided by converting the audio model (via audio conversion) and the video model from the reference individual to the target person. Emotional effects are achieved by recognizing emotions in the audio and / or video data of the reference individual.
The synthesis of the direct TTS speech model will be described.
When the target audio is produced using the disclosed TTS technology, a dictionary that maps words and phonemes to a time series of multipliers is used to obtain synchronized video synthesis. As mentioned above, the alignment transformation matrix is used between the reference and target video models. In one embodiment, if the target person is a reference individual and the dictionary is based on the reference individual, then the alignment transformation matrix is unnecessary and the dictionary contains audio and video sequences of the target person. Used to adjust directly. In other embodiments, it does not have a dictionary based on the target person. The multiplier is calculated based on the dictionary of the reference individual, and this multiplier is converted into the multiplier for the target person by using the alignment conversion matrix by one calculation.
FIG. 10 shows an example of the process of synthesizing a video string and an audio string. The letter 1010 is decomposed into words, phonemes, and utterances (1020). Each word, or phoneme, or utterance has a duration of 1030. Each word, or phoneme, or utterance matches the entry in the dictionary (1040). The registration includes a time series in a multiplier or VQ center (or VQ index). The process checks to see if the duration of spoken words or phonemes produced by the TTS system matches the duration of the corresponding visual movement created by the dictionary. If the durations do not match, the situation is corrected by the first and last timestamps created by the TTS system (1050). In this way, synchronization between the audio sequence generated by the TTS and the video sequence generated by the dictionary is achieved.
Thus, by applying the reference to the target alignment transformation (1070) to generate each frame of the target person's video, a properly synchronized frame of the video sequence 1060 for the reference individual can be obtained. Will be generated.
The matching of the generated image with the background will be described.
In the above, we focused on generating proper facial and mouth movements for the target person. Other parts of the body (especially paper and neck, shoulders) need to be generated for a complete picture of the target person. The "background" in this disclosure includes two areas. The first area is the target human body that is not generated by the video model, and the second area is a landscape different from the target body.
The work of embedding a composite video in the background is a layering procedure. The background is covered by the body portion, and the composite video portion fills the rest of each frame. In one embodiment, there are no restrictions on the selection of the background for video composition. For example, the user can select the desired background from the menu to achieve a particular effect.
Since the body must naturally fit the target person's face, there are further restrictions on the target person's body that is not generated by the video model. There are multiple techniques for addressing this embedded portion, disclosed in the following paragraphs, which are combined with each other.
An embedding method that minimizes the boundary fit error will be described.
FIG. 11 shows an example of the process of embedding a composite video in the background while minimizing the boundary matching error. Each frame 1110 and the background image of the generated target image (here, the image part of the image includes only the synthesized face part of the target person) are matched (stitched) with each other. The coordinates of the boundary points of the face are calculated and recorded. A search is performed on the background image for the optimum area of the background image where the difference between the boundary of the optimum area of the background image and the boundary point of the composite image is the smallest. Once the optimum background region is specified, the optimum frame 1120 of the background image is determined in which the boundary error of the first frame 1130 of the combined target image is minimized. Subsequently, the boundary point of the synthesized target image is moved to the background image. The shape multiplier is adjusted and resynthesized based on the coordinates of the internal (target) and external (background) points for a realistic composition of the target person. This process is repeated until the boundary point is within a certain margin of error for the background boundary point. The synthesized face is embedded inside the non-synthetic part (1140). Since the frame from the background image is selected to minimize the embedding error as described above, the movement of the face (particularly the position of the mouth) is minimized.
Continue to the composited frame of the next video sequence. The same boundary error is calculated for the next composite frame, the previously used footage, and the previous frame of the background footage. Of these three frames, find the video that minimizes the boundary error and repeat the iterative process outlined above to embed the second frame. This process is repeated for each frame of the composite video (1150).
Embedding by region division will be described.
As mentioned in the previous paragraph, separate models are used for the upper and lower parts of the face. FIG. 12 shows an example of the process of embedding a composite video in the background using a two-region model. In this model, the upper and lower parts of the face are fitted to the existing image of the target person, i.e. the background image (1210).
The upper boundary point (such as on the forehead) is relatively rigid (not moving) and is aligned with the top to the boundary using the simple stiffness transform 1220 (which changes direction, including scaling). By moving all the points in the part, the upper face is embedded in the background (1230).
Since the boundary point of the lower region (the chin point is on the boundary in relation to speaking) is non-rigid, the lower region is not embedded in the background in the same manner as the upper region. However, some information comes from the synthesized boundaries. The lower area is sized appropriately to embed the lower face in the background (1240). This provides scaling parameters and facilitates alignment of the lower region to the upper region.
The upper region and the lower region are connected to each other in the following manner. The connection between the lower and upper regions is performed so that the two regions have at least three points in common. This commonality determines how to move, rotate, and resize the lower region to connect to the upper region (1250). The lower regions are aligned according to commonalities (1260), and the upper and lower regions are combined to form an embedded entire face (1270).
The area of interest will be described.
The background is divided into multiple regions of interest (ROIs). For example, areas such as the neck and shoulders are included in the area of interest. The boundaries of the composite video sequence are tracked. The frame containing the maximum fit between the composite video sequence and the region of interest, including the neck and shoulders, is selected as the basis for embedding the composite video sequence in the background video. Techniques that utilize the domain of interest are disclosed in detail in US patent application 13 / 334,726, which is incorporated herein by reference.
The techniques introduced herein are performed, for example, by a programmable circuit (eg, one or more microprocessors) programmed in software and / or firmware, or for a specific purpose as a whole. Wiring-connected circuits for specific purposes consist of, for example, one or more application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), and the like.
The software or firmware used to perform the techniques introduced herein is recorded on a computer-readable storage medium and executed by one or more general purpose or specific purpose programmable microprocessors. The term "computer-readable storage medium" as used herein is connected by a machine (eg, a computer, network equipment, mobile phone, personal digital assistant (PDA), manufacturing tool, device with one or more processors, etc.). Includes any device that records information in a possible form. For example, the storage medium to which a computer can be connected is a recordable or non-recordable medium (for example, read-only memory (ROM), random access memory (RAM), magnetic disk recording medium, optical recording medium, flash memory, disk drive, etc. ) Etc. are included.
As used herein, the term "logic" includes, for example, programmable circuits programmed in specific software and / or firmware, wiring-connected circuits for specific purposes, or a combination thereof.
In addition to the above examples, various improvements and modifications of the invention are possible without departing from the spirit of the invention. Therefore, the present invention is not limited to the above disclosure, and the claims are construed to include the gist and the whole scope of the present invention.
In one embodiment, it is about a method. In this method, in order to model a step of inputting a character string into a processing device and a visual and audible expression of a person's emotions, the processing device uses the processing device to generate a video string of the person based on the character string. The step of generating includes the step of using the voice model of the voice of the person to generate the voice portion of the video string.
In a related embodiment, the processing device is a mobile device, the character string is input from a second mobile device via short message service (SMS), and the step of generating a human video string is the step. This includes generating a video sequence of a person by the mobile device based on the shared information recorded in the mobile device and the second mobile device.
In another related embodiment, the string has a group of words that includes at least one word, and the video string is generated so that the person appears to speak a word in the video string.
In another related embodiment, the string has characters that represent vocalizations, and the video string is generated so that the person appears to speak a word in the video string.
In another related embodiment, the character string has a word and an index of the word, and the index is the image when the person appears to speak the word in the image string. The emotional expression of the person is displayed simultaneously in the column, the index is a defined index group, and each index of the defined index group is related to a different emotional expression.
In another related embodiment, the step of generating a video sequence is to model a visual and audible emotional expression of the person based on the string and the person's prior knowledge. Includes generating a human video sequence with a processing device.
In other related embodiments, the prior knowledge includes a photograph or video of a person.
In another related embodiment, the step of generating the video sequence includes a step of mapping the words of the string to the features of the person's face and a step of rendering the features of the person's face on the background. Including.
In other related embodiments, the word is mapped to the facial features based on one or more indicators for the word, which indicators are such that the person utters the word in the video sequence. When it looks like this, the emotional expression of the person is displayed at the same time in the video sequence.
In other related embodiments, the facial features include general facial features that apply to more than one person.
In other related embodiments, the facial features include specific facial features that apply to the particular person.
In another related embodiment, the step of generating a sequence of images comprises generating a gesture of the person's body that matches the facial features of the person.
In another related embodiment, the step of generating a video sequence uses an audio model based on the person's audio to generate an audio string that represents the person's audio based on the words in the string. Including that.
In another related embodiment, receiving the string comprises receiving the string in real time, and the step of generating the video string models a visual and audible emotional expression of the person. Therefore, it includes a step of generating a person's video string in real time based on the character string, and the step includes using an audio model of the person's voice to generate an audio portion of the video string.
In other related embodiments, it is about other methods. This method includes a step of inputting a character string into a processing device and a step of generating a video string of the person based on the character string by the processing device in order to model a visual expression of a person's emotions. In order to model the steps in which the face of each frame of the video string is represented by the combination of a plurality of inferred images of the person, and the audible, emotional expression of the person, the voice of the person. Using the audio model, the step of generating the audio string of the person based on the character string by the processing device, and the video of the person by combining the video string and the audio string using the processing device. A step of generating a column, which includes a step of synchronizing the video string and the audio string based on the character string.
In another related embodiment, the face of each frame of the video sequence is represented by a linear combination of the plurality of guess images of the person, and each guess image in the plurality of guess images of the person is the average of the person. Corresponds to the deviation from the image.
In another related embodiment, the step of generating the person's video sequence based on the string comprises dividing each frame of the video string into two or more regions, the at least one said region. , Expressed by combining the guessed images of the person.
In another related embodiment, the voice model of the person's voice comprises a plurality of voice features made from the person's voice sample, and each voice feature of the plurality of voice features corresponds to a character. To do.
In other related embodiments, each voice feature in the plurality of voice features corresponds to a word, or phoneme, utterance.
In another related embodiment, the voice model of the person's voice comprises a plurality of voice features made from the person's voice sample, a second person's voice following the string, and a waveform of the person's voice. And the match (correspondence) with the voice waveform of the second person, and the voice feature of the person is based on the match (correspondence) between the voice waveform of the person and the voice waveform of the second person. Then, it is mapped to the voice of the second person.
In another related embodiment, the voice model of the person's voice includes a plurality of voice features made from the person's voice sample and the voice produced by the model that converts the characters into voice according to the string. , The person's voice feature includes the match (correspondence) between the person's voice waveform and the voice waveform of the model that converts characters into voice, and the person's voice feature is a model that converts the person's voice waveform and characters into voice. It is mapped to the voice of the model based on the match (correspondence) with the voice waveform of.
In other related embodiments, it is about other methods. This method is a step of generating a character string, which is generated using a voice model based on the person's voice in order to visually and audibly represent the range of the person's emotions. A step configured to express one or more words uttered by a person in a video string and a step of identifying an index related to the word in the character string, and the index is a specified index. One of the groups, a step in which each index is configured to show a different emotional expression of the person, a step in which the index is incorporated into the character string, and a device configured to generate the video string. It includes a step of transmitting the character string.
In other related embodiments, the step of identifying an indicator includes selecting one item from a list of items related to a word in the string, where each item in the list is the person's emotions. It is an index showing the expression.
In other related embodiments, the step of identifying an indicator comprises inserting a markup language string used for a word in the string, the markup language string being a default markup language character. It consists of a group of columns, and each markup language string in the group is an index showing a person's emotional expression.
In another related embodiment, the step of identifying an indicator uses an automatic speech recognition (ASR) device and is based on the speech string of a speaker who speaks the words in the string. Includes steps to identify indicators related to.
In other related embodiments, the speaker is a different person than the person.
In other related embodiments, it is about other methods. This method is based on a step of storing information of a plurality of non-person items in a processing device and prior information of the non-individual items of the processing device. A step of generating a video sequence for an item, including a step in which each of the non-individual items is configured to be independently manageable.
In other related embodiments, the non-individual items are constrained in relation to other elements in the video sequence.
14 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9959657B2 | Cited by | United States of America | Applicant |
| JP2014146340A | Cited by | Japan | Search report |
| JP2024532728A | Cited by | Japan | Search report |
| JP2014146339A | Cited by | Japan | Search report |
| JP2003216173A | Cites | Japan | Search report |
| JP2007279776A | Cites | Japan | Search report |
| JP2010277588A | Cites | Japan | Search report |
16 members in 6 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 201161483571 | United States of America | P | |
| 201161483571 | United States of America | P | |
| 61483571 | United States of America | – | |
| 2012036679 | United States of America | W | |
| 2012036679 | United States of America | W | |
| 2011483571 | – | – | – |
| 2012036679 | – | – | – |
| US201161483571P | – | – | – |
| WO2012US36679 | – | – | – |
Members16
| Document | Office | Kind | |
|---|---|---|---|
| WO2012088403A2 | World Intellectual Property Organization (WIPO) | A2 | |
| TW201236444A | Taiwan Province of China | A | |
| WO2012088403A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2012154618A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2012327243A1 | United States of America | A1 | |
| WO2012154618A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US2013124206A1 | United States of America | A1 | |
| EP2705515A2 | European Patent Office (EPO) | A2 | |
| CN103650002A | China | A | |
| JP2014519082AThis record | Japan | A | |
| EP2705515A4 | European Patent Office (EPO) | A4 | |
| US9082400B2 | United States of America | B2 | |
| JP6019108B2 | Japan | B2 | |
| CN103650002B | China | B | |
| CN108090940A | China | A | |
| US10375534B2 | United States of America | B2 |
17 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written submission of copy of amendment under article 19 pctJAPANESE INTERMEDIATE CODE: A524A524 | A524 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2014519082
- Publication, DOCDB
- 2014519082
- Publication, EPODOC
- JP2014519082
- Application
- 2014509502
- Application, DOCDB
- 2014509502
- Application, EPODOC
- JP20140509502
Titles2
- Japanese
- 文字に基づく映像生成
- English
- Video generation based on characters
Classification
- CPC, 5
- G06T13/40
- G10L13/08
- G10L13/10
- H04M1/72436
- H04M1/72439
- IPC, 5
- G06T13 40
- G10L13 00
- G10L15 10
- H04M1 72436
- H04M1 72439
Designated states5
- Regional, 4
- Zimbabwe
- Turkmenistan
- Türkiye
- Togo
- National, 1
- South Africa