System and method for providing audio for a requested note using a render cache.
Abstract
A method for providing audio data corresponding to a requested musical note is disclosed, the method comprising: (a) providing a render cache having a plurality of cache entries, each of the cache entries corresponding to a different note; (b) receiving a request for a first note from a client; (c) identifying a first cache entry corresponding to the first note; (d) determining that a first audio segment corresponding to the first cache entry is not available; (e) identifying a second audio segment corresponding to a near-hit cache entry in the render cache; and (f) processing the second audio segment into a third audio segment that is substantially similar to the first audio segment.

Term
5.8 yearsleft in the term
Expires 30 July 2032.
- Priority
- Filed
- Granted
- Today
- Expires
34 claims: 3 independent, 31 dependent
- 1Claims Reivindicaciones 1. A method of providing audio data corresponding to a requested musical note, the method comprising:1. Un método para proporcionar datos de audio correspondientes a una nota musical solicitada, comprendiendo el método: receive a request that includes a first note from a customer;recibir una solicitud que incluye una primera nota de un cliente;determinar, a través de uno o más procesadores, una primera entrada de caché en una memoria caché de procesamiento que incluye una pluralidad de entradas, la primera entrada de caché correspondiente a la primera nota;determining, through one or more processors, a first cache entry in a processing cache including a plurality of entries, the first cache entry corresponding to the first note;determinar, a través de los uno o más procesadores, que un primer segmento de audio correspondiente a la primera entrada de caché está ausente;determining, through the one or more processors, that a first audio segment corresponding to the first cache entry is absent;en respuesta a la determinación de que el primer segmento de audio está ausente, la determinación, a través de los uno o más procesadores, que una entrada de caché de resultado-cercano está disponible en la memoria caché de procesamiento;in response to determining that the first audio segment is absent, determining, through the one or more processors, that a near-result cache entry is available in the render cache;en respuesta a la determinación que la entrada de caché de resultado-cercano está disponible, recuperar un segundo segmento de audio correspondiente a la entrada de caché de resultado-cercano en la memoria caché de procesamiento;y in response to determining that the near-result cache entry is available, retrieving a second audio segment corresponding to the near-result cache entry in the processing cache;Y 208 Τ Μ Ρ 1 208 Τ Μ Ρ 1 INSTITUTO MEXICANO MEXICAN INSTITUTE DE LA PíOFIEDAD O-Biw-jSi OF THE PIOFTY O-Biw-jSi INDUSTRIAL 7 * - generate a third audio segment from the second audio segment, the third audio segment being substantially similar to the first audio segment. INDUSTRIAL 7* — generar un tercer segmento de audio desde el segundo segmento de audio, siendo el tercer segmento de audio sustancialmente similar al primer segmento de audio.
- 27A method to provide audio for a requested note, comprising the method:27. Un método para proporcionar audio para una nota solicitada, comprendiendo el método: 213 213 receive a first audio request, including a first note from a first customer, the first audio request, including identifying information of a unique note ID of the first note, a unique track ID of a track associated with the first note , and a start time of the first note within the track associated with the first note;recibir una primera petición de audio, incluyendo una primera nota de un primer cliente, la primera petición de audio, incluyendo información de identificación de un ID de nota única de la primera nota, una ID de pista única de una pista asociada con la primera nota, y un tiempo de inicio de la primera nota dentro de la pista asociada con la primera nota ;determinar, a través de uno o más procesadores, una primera cola de pista basada en el ID de pista única, la primera cola de la pista siendo una de una pluralidad de colas de pista;determining, via one or more processors, a first track queue based on the unique track ID, the first track queue being one of a plurality of track queues;determinar, a través de los uno o más procesadores, que la primera cola de pista incluye una solicitud de audio recibida previamente que incluye el ID de nota única;determining, through the one or more processors, that the first track queue includes a previously received audio request that includes the unique note ID;en respuesta a la determinación de que la primera cola de pista incluye la solicitud de audio recibida anteriormente que incluye el ID de nota única, remover la solicitud de audio recibida previamente de la cola de pista;in response to determining that the first track queue includes the previously received audio request that includes the unique note ID, removing the previously received audio request from the track queue;basado en el tiempo de inicio de la primera nota, determinar, a través de los uno o más procesadores, una primera posición de la primera solicitud de audio dentro de la primera cola de pista con respecto a por lo menos una solicitud de audio adicional recibida previamente;y based on the start time of the first note, determining, through the one or more processors, a first position of the first audio request within the first track queue with respect to at least one additional received audio request previously;Y 214 214 IMPI IMPI INSTITUTO MBXICANí INSTITUTO MBXICANí DE LA PaOrilDAD INDUSTRIAL OF INDUSTRIAL ROLE add the first audio request to the first track queue for processing at the first position. adicionar la primera solicitud de audio a la primera cola de pista para el procesamiento en la primera posición.
- 28A process to provide audio for an unsolicited; understanding the procedure:28. Un proceso para proporcionar audio para un no solicitado;comprendiendo el procedimiento: receiving a first audio request from a client, the first audio request including a first note on an audio track;the client being configured to repeatedly play a live loop that includes at least a portion of the audio track;the first audio request including timing information regarding the live loop;recibir una primera solicitud de audio de un cliente, la primera solicitud de audio incluyendo una primera nota en una pista de audio;el cliente estando configurado para reproducir repetidamente un bucle en vivo que incluye al menos una porción de la pista de audio;la primera solicitud de audio que incluye información de temporización relativa al bucle en vivo;determinar, a través de uno o más procesadores, un primer tiempo para servicio para la primera solicitud de audio, el primer tiempo para servicio siendo indicativo de un tiempo hasta que una próxima instancia en la que se reproducirá la primera nota dentro del bucle en vivo;determine, through one or more processors, a first service time for the first audio request, the first service time being indicative of a time until a next instance in which the first note will be played within the live loop ;comparar, a través de los uno o más procesadores, el primer tiempo para dar servicio a los tiempos anteriores a servicio asociado con al menos una solicitud de audio’ previa;comparing, through the one or more processors, the first service time to the previous service times associated with at least one previous audio request ';en respuesta a la comparación entre el primer tiempo para servicio y los tiempos anteriores para servicio, in response to the comparison between the first service time and previous service times, 215 ,. . ,,, UDUÍTIUAL determine, a first position of the first audio request within a queue;and adding the first audio request to the queue for processing in the first position. 215 , . . , , , UDUÍTIUAL determinar, una primera posición de la primera solicitud de audio dentro de una cola;y la adición de la primera solicitud de audio a la cola para su procesamiento en la primera posición.
Independent claims3
897 paragraphs in 38 sections, as filed
(54) Title: SYSTEM AND METHOD TO PROVIDE AUDIO FOR A REQUIRED NOTE USING A PROCESSING CACHE MEMORY.
(54) Title: SYSTEM AND METHOD FOR PROVIDING AUDIO FOR A REQUESTED NOTE USING A RENDER CACHE.
(57) Summary
A method for providing audio data corresponding to a requested musical note is disclosed, the method comprising: (a) providing a processing cache having a plurality of cache entries, each of the cache entries corresponding to a different note ; (b) receive a request for a first note from a customer; (c) identifying a first cache entry corresponding to the first note; (d) determining that a first audio segment corresponding to the first cache entry is not available; (e) identifying a second audio segment corresponding to a close hit cache entry in the processing cache; and (f) processing the second audio segment into a third audio segment that is substantially similar to the first audio segment.
(57) Abstract
A method for providing audio data corresponding to a requested musical note is disclosed, the method comprising: (a) providing a render cache having a plurality of cache entries, each of the cache entries corresponding to a different note; (b) receiving a request for a first note from a Client; (c) identifying a first cache entry corresponding to the first note; (d) determining that a first audio segment corresponding to the first cache entry is not available; (e) identifying a second audio segment corresponding to a near-hit cache entry in the render cache; and (f) Processing the second audio segment into a third audio segment that is substantially similar to the first audio segment.
Mexican Institute of Industrial Property
<img file="MX345588B_D0001.tif" />
PATENT TITLE NO. 345588
Headlines)
MUSIC MASTERMIND, INC.
Home:
22315 Mulholland Highway, Calabasas, California, 91302, USA
Denomination:
SYSTEM AND METHOD FOR PROVIDING AUDIO FOR A REQUIRED NOTE USING A PROCESSING CACHE MEMORY.
Classification:
lnt.CI.8: A63H5 / 00; G01H7 / 00; G06F3 / 0481; G09B15 / 00
Inventor (s):
REZA RASSOOL: DARREN WARNER; MATT SERLETIC
REQUEST
Number:
International filing date!
MX / a / 2014/001194 of July 2012
PRIORITY
Country:
Date:
Number;
US July 2011
13/194,806
Validity: Twenty years Expiration date: July 30, 2032
The reference patent is granted based on articles 1<sup>or</sup>, 2<sup>or</sup> fraction V, 6<sup>or</sup> fracc in III, and 59 of the Industrial Property Law
In accordance with article 23 of the Industrial Property Law, this patent is valid for ten years, renewable, counted from the filing date of the international application and will be subject to payment to keep it in force. rights.
i;
^ ulen subscribes this title k> does with base on the provisions of the article 6 * fractegnes III and 7 ° bis 2 of the Industrial Property Law (Official Diteio of the Federation (Γ OF.) 2<sup>7</sup>/ 0ñ <l991, reform OSWtOM, 10/25/1996, 12/26/19 ^, 05/17/1999, 01/26/2004, 06/16/2005,25/01/20Q6, 05/06/ 2009.06 / 01/2011 1 406/2010, 26/00/2010, 01/27/2012 v 04/09/2012); items 1® 3<sup>or</sup> fraction V ¿ciso a), 4th and 12th paragraphs i and Jl of the Regulations of the Instituto Medaño of such Industnal Property (DOF 14/12 / 199® amended on ¢ 1/07/2002, 15/07/2004,> 8 / 07 / 20® »and 7/09/2007); items 1<sup>or</sup>, 3<sup>or</sup>, 4<sup>or</sup>, 5th section V subsection a) 16 sections I and III and 30 of the Organic Statute of the Mexican Institute of Industrial Property! (DO F. 12/27/1999. Amended if 1W10 / 2002, as / mowazotM and 08/13/2007); 1, 3 and 5 subsection a) of the Agreement that delegates powers to the Deputy General Directors, Coordinator, Divisional Directors, Heads of Regional Offices, Divisional Deputy Directors, Departmental Coordinators and other subordinates of the Mexican Institute of Industrial Property. (DOF 12/15/1999, amended on 02/04/2000, 07/29/2004, 08/04/2004 and 09/13/2007).
<img file="MX345588B_D0002.tif" />
You
Coi
X;
No 550. Floor r Santa María Tepepa;
Issue Date: February 7, 2017
THE DIVISIONAL DIRECTOR OF PATENTS
<img file="MX345588B_D0003.tif" />
NAHANNY CANAL REYES
<img file="MX345588B_D0004.tif" />
SYSTEM AND METHOD OF PROVIDING AUDIO FOR A
<img file="MX345588B_D0005.tif" />
<img file="MX345588B_D0006.tif" />
USING A MEMORY PROCESSING CACHE
This application claims priority from US Provisional Patent Application No. 61 / 182,982, filed 1/6/2009; US Provisional Patent Application No. 61 / 248,238, filed 2/10/2009; US Provisional Patent Application No. 61 / 266,472, filed 3/12/2009; US Provisional Patent Application 10 Nos. 12/791/792; 12 / 791,798; 12 / 791,803 and 12 / 791,807 all filed on June 1, 2010.
Technical field
The present invention relates generally to music creation, and more particularly to a system and method for providing audio for a required note using a processing cache.
Background
Music is a credibly well-known form of human expression. However, a person's first-hand appreciation for this artistic creation can be derived in different ways. Often times, the person can enjoy music more easily by listening to the creations of others rather than generating it himself or herself. For many people, the ability to recognize a beautiful musical composition is innate, while the ability to manually create an appropriate collection of notes remains out of reach. A person's ability to create new music can be inhibited by time, money, and / or the skill necessary to learn an instrument well enough to accurately reproduce a tune at will. For many people, their own imaginations can be the source of new music, 10 but their ability to hum or sing this same tune limits the scope for which their tunes can be formally retained and recreated for the enjoyment of others.
Recording a performance session of a musician can also be a laborious process. Multiple takes of the same material are recorded and carefully scrutinized until a single take can be pieced together with all imperfections corrected. A good take often requires a talented artist under the direction of another to adjust its execution accordingly. In the case of an unprofessional 20 recording, the best take is often the result of chance and therefore cannot be repeated. Most of the time, amateur performers produce takes with both good and bad parts. The recording process would be much easier and more fun if a song could be constructed without having to
,. ,. . ΙΜΡΙβ meticulously analyze each portion of each iInkfaaAMo
It is therefore with respect to these and other considerations that the present invention has been made.
Also, the music that a person wants to create can be complex. For example, an imagined melody may have more than one instrument, which must be played at the same time with other instruments in a potential arrangement. This complexity further adds to the time, skill and / or money required by a single person to generate a desired combination of sounds. The physical configuration of many musical instruments also requires a person's complete physical attention to manually generate notes, and requires additional personnel to play the additional parts of a desired melody. Additionally, extra review and direction may be needed to ensure the proper interaction of the various instruments and elements involved in a desired melody.
Even for people who already enjoy creating their own music, those listeners may lack the kind of knowledge that allows for proper music creation and composition. Consequently, the music created may contain notes that are not within the same key or chord. In most musical styles, the presence of notes out of tune or chord, often called dissonant notes, causes the music to be discordant and unpleasant. Consequently, due to their lack and training, listeners frequently create music that sounds undesirable and unprofessional.
For some people, artistic inspiration is not subject to the same location and time constraints that are typically associated with the generation and recording of new music. For example, a person may not be in a production studio with an instrument available to play in hand when an idea for a new melody materializes. After the moment of inspiration passes, the person may be unable to recall the full extent of the original melody, resulting in a loss of artistic effort. Additionally, the person can become frustrated with the time and effort put into recreating no more than a bottom and an incomplete version of their initial musical reveal.
Computer program (software) tools for professional music editing and composition are generally available today. However, these tools project an intimidating barrier to be crossed by a novice user. Such complex user interfaces can quickly sap the enthusiasm of any beginner trying to venture their way on a sudden artistic desire. Having to rely on a set of professional audio servers too
IMPIO ττπντο MSXICANi limits the style of the creative mobile, wanting to create üTra. melody on the go. --------------- What is needed is a system and method of music creation that can easily interact with the basic ability of the user, and yet allow the creation of music that is so complex as the user's expectations and imagination. There is also an associated need to facilitate the creation of music free of notes that are dissonant. Furthermore, there is a need in the art for a music creation system that can generate a compiled music track by adding portions of multiple takes based on automated selection criteria. It is also desirable that such a system can also be implemented in a way that is not limited by the location of a user when inspiration appears, thus allowing the capture of the first expressions of a new musical composition.
There is an associated need in the art for a system and method that can create a compiled track from multiple takes by automatically evaluating the quality of the previously recorded tracks and selecting the best of the previously recorded tracks, these being recorded by means of a electronic creation system.
It is also desirable to implement a system and method for creating music that is cloud-based whereby intensive processing functions are
NSTmKO MHXICAHO implemented by a remote server from a di ^] b ^ £ ^ & vo ^ client. However, because the creation of -di-gita-l · of .....— music depends on large amounts of data, such settings are generally limited by several factors. For the provider, processing, storing and servicing such large amounts of data can be overwhelming unless the central processor is extremely powerful and therefore expensive from a cost and latency standpoint. Given the current costs of storing and sending data, transmitting data from a processing server to a client can quickly become cost prohibitive and can add unwanted latency. From a customer perspective, bandwidth limitations can also lead to significant latency issues, detracting from the user experience. Thus, there is also a need in the art for a system that can address and overcome these disadvantages.
Brief description of the drawings
Non-limited and non-exhaustive modes are described with reference to the following drawings. In the drawings, the same reference numbers refer to like parts throughout the various figures unless otherwise specified.
i<sup>Μ</sup>
For a better understanding of this document, a reference will be made to the following detailed description, which should be read in conjunction with the accompanying drawings, where:
Figs. 1A, IB and 1C illustrate various embodiments of a system in which aspects of the invention can be practiced;
Figure 2 is a block diagram of one embodiment of potential components of the audio converter (140) of the system of Figure 1;
Figure 3 illustrates an exemplary embodiment of a progression for a musical compilation;
Figure 4 is a block diagram of one embodiment of potential components of the 15-track partitioner (204) of the system of Figure 2;
Figure 5 is an example frequency spectrum diagram illustrating the frequency distribution of an audio input with a fundamental frequency and multiple harmonics;
Figure 6 is an example plot of pitch versus time illustrating the pitch of a human voice changing between a first and a second pitch and settling near the second pitch;
IMPI O?
'ismvro mexican
Figure 7 is an example ^ 'modality of ^^ na ^ morphology plotted as pitch events over time, each having a discrete duration;
Figure 8 is a block diagram illustrating the contents of a data file in one embodiment of the invention;
Figure 9 is a flow chart illustrating one embodiment of a method for generating music tracks within a continuous loop recording session;
Figs. 10, 10A and 10B together form an illustration of a potential user interface for generating music tracks within a continuous loop recording session;
Figure 11 is an illustration of a potential user interface for calibrating a recording session;
Figs. 12A, 12B, and 12C together illustrate a potential second user interface associated with generating music tracks within a recording session with continuous looping over three separate periods of time;
Figs. 13Ά, 13B, and 13C together illustrate a potential use of the user interface to modify an input music track into the system using the user interface of Figure 12;
'• ΤΠΤΙΓΓΟ MWGCAN ·. · ..
Figs. 14Α, 14Β and 14C together illus<sup>:</sup>ti? arí · —uiiaSB-potential user interface to create a track from. ritme --- in three separate periods of time;
Figure 15 is a block diagram of an embodiment 5 of potential components of the MTAC module (144) of the system of Figure 1;
Figure 16 is a flow chart illustrating a potential process for determining the musical key reflected by one or more notes of the audio input;
Figure 16A illustrates an interval profile matrix that can be used to better determine the reinforcement.
Figs. 16B and 16C illustrate Major and Minor Key Interval Profile Matrices, respectively, which are used in association with the interval profile matrix to provide a determination of a preferred hue.
Figs. 17, 17A and 17B together form a flow chart illustrating a potential process for rating a portion of a music track based on a chord sequence constraint;
Figure 18 illustrates one embodiment of a process for determining the centroid of a morphology;
Figure 19 illustrates stepped responses of a harmonic oscillator over time having a
ΙΜΡΠ attenuated response, an over-attenuated response '^^ ™ ^^ or ^^ under-attenuated response;
Figure 20 illustrates a logic flow diagram showing one modality for rating a portion of a music entry;
Figure 21 illustrates a logic flow diagram for one embodiment of a process for composing the best track from multiple recorded tracks;
Figure 22 illustrates one embodiment of an example audio waveform 10 and a graphical representation of a note showing the difference of the current pitch to the ideal pitch;
Figure 23 illustrates an embodiment of a new track constructed from previously recorded track partitions 15;
Figure 24 illustrates a data flow diagram showing one embodiment of a process for harmonizing an input musical accompaniment with an input main music;
Figure 25 illustrates a data flow diagram of the processes executed by the Notes Transformation Module of figure 24;
Figure 26 illustrates an exemplary embodiment of a super keyboard;
JMPI
Figs. 27A-B illustrate two modalities of ej
<img file="MX345588B_D0007.tif" />
a circle of fifths; ———, -
Figure 28 illustrates an exemplary embodiment of a network configuration in which the present invention can be practiced;
Figure 2-9 illustrates a block diagram of a device that supports the processes discussed herein;
Figure 30 illustrates one embodiment of a music network device;
Figure 31 illustrates a potential embodiment of a first interface in a gaming environment;
Figure 32 illustrates a potential embodiment of an interface for creating one or more main vocal or instrumental tracks in the game environment of Figure 31;
Figure 33 illustrates a potential embodiment of an interface for creating one or more drum tracks in the game environment of Figure 31;
Figs. 34A-C illustrate potential modalities of an interface for creating one or more backing tracks in the game environment of Figure 31;
Figure 35 illustrates a potential embodiment of a graphical interface that represents the chord progression being played in accompaniment to the main music;
<img file="MX345588B_D0008.tif" />
<img file="MX345588B_D0009.tif" />
Figure 36 illustrates a 'ϊΝύυπτυΑΐ mode selecting between different sections of a music compilation in the game environment of Figure 31;
Figs. 37A and 37B illustrate potential embodiments of a file structure associated with music assets that can be used in conjunction with the gaming environment of Figures 31-36;
Figure 38 illustrates one embodiment of a cache memory depicted in accordance with the present invention;
Figure 39 illustrates one embodiment of a logic flow diagram showing one mode for obtaining audio for a requested note in accordance with the present invention;
Figure 40 illustrates one embodiment of a flow chart for implementing the cache control process 15 of Figure 39 in accordance with the present invention;
Figure 41 illustrates one embodiment of an architecture for implementing a processing cache in accordance with the present invention; Y
Figure 42 illustrates a second embodiment of an architecture for implementing a processing cache in accordance with the present invention.
Figure 43 illustrates one embodiment of a signal diagram illustrating communications between a client, a server, and an edge cache in accordance with the present invention.
I Ad PI • NST1TUTO MEXICANO
OF THE PROHEDA »
Figure 44 illustrates a second mode<sup>IN</sup>™ -Sytt ufft—— signal diagram illustrating communications between client UI1, a server, and an edge cache in accordance with one embodiment of the present invention.
Figure 45 illustrates one embodiment of a first process for optimizing an audio request processing queue in accordance with the present invention.
Figure 46 illustrates one embodiment of a second process for optimizing an audio request processing queue in accordance with the present invention.
Figure 47 illustrates an embodiment of a third process for optimizing an audio request processing queue in accordance with the present invention.
Figure 48 illustrates an exemplary embodiment of a live loop performance in accordance with one embodiment of the present invention.
Figure 49 illustrates one embodiment of a series of effects that can be applied to a compilation of music according to the present invention.
Figure 50 illustrates one embodiment of a series of effects corresponding to the role of the musician that can be applied to a track of an instrument according to the present invention.
Figure 51 illustrates a modality of a series of effects of the role of the producer that can be applied to
<img file="MX345588B_D0010.tif" />
Ι.ΜΡΙ a track of an instrument according to, 'iÑnusTWAi invention.
Figure 52 illustrates one embodiment of a series of effects corresponding to the producer role that can be applied to a compilation track according to the present invention.
Detailed description
The present invention will be described more fully hereinafter with reference to the accompanying drawings, which form a part of the present document, and show, by way of illustration, specific exemplary embodiments by which the invention may be practiced. . This invention can be made in many different ways and should not be construed as limited to the embodiments set forth herein; rather; These embodiments are provided so that this description is thorough and complete, and fully expresses the scope of the invention for those skilled in the art. Among other things, the present invention can be carried out as methods or devices. Accordingly, the present invention may take the form of a physical equipment (hardware) mode as a whole, a software-only mode, or a mode combining software and hardware aspects. The following detailed description is therefore limiting.
should not be taken
<img file="MX345588B_D0011.tif" />
Definitions *
Throughout the description and claims, the following terms take on the explicitly associated meanings herein, unless the context clearly dictates otherwise. The phrase in an embodiment as used herein does not necessarily refer to the same embodiment, although it may be. Furthermore, the phrase in another embodiment as used herein does not necessarily refer to a different embodiment, although it may be. Thus, as described below, different embodiments of the invention can be easily combined, without departing from the scope or spirit of the invention.
Additionally, as used herein, the term or is an inclusive operator or, and is equivalent to the term and / or, unless the context clearly dictates otherwise. The term based on is not exclusive and is allowed to be based on additional factors not described, unless the context clearly dictates otherwise. Additionally, throughout the description, the meaning of a, an, and the includes the plural references. The meaning of en includes in and includes plural references. The meaning of en
ΪΜΡΙ ^
ÍNrJUST ^ IAL about.
The term "musical input" as used herein, refers to any input signal that contains musical and / or control information transmitted by any or a variety of media, including, but not limited to over the air, microphones, mechanisms in line, or the like. Musical inputs are not limited to input frequency signals that can be heard by the human ear, and may include other frequencies outside of those that can be perceived by the human ear, or in a way that is not easily audible to the ear. human. Furthermore, the use of the term musical is not intended to express an inherent requirement for a pulse, rhythm or the like. Thus, for example, a musical input may include different inputs such as a tapping, including a single hit, clicking (clicking), human inputs (such as voice (for example do, re, mi), percussed inputs (for example ka , cha, da-da), or the like) as well as indirect inputs through an instrument or other amplitude and / or frequency generating mechanism through a means of transport including, but not limited to, a microphone input , an online entry, a MIDI input, a file that has useful signal information for expressing
IMPIOS a musical entry, or other entries that uíiavrsOj
INDUSTRIAL U-Saj ^ transported signal can be converted into music.
The term musical tonality as used herein is a group of musical notes that are harmonious. The hues are usually major or minor. Musicians frequently speak of a musical composition that is in the key of C major, for example, which implies that a musical piece harmonically centered on the note C and making use of the major scale whose first note or tonic is C. A major scale is an eight-note progression consisting of the perfect, major semitones (for example CDEFGAB or do re mi fa g la si). With respect to a piano, for example, the middle C (sometimes called C4) has a frequency of 261.626 Hz, while that of
D4 is 293,665 Hz; E4 is 329,628 Hz; F4 is 349,228 Hz; G4 is 391,995 Hz; A4 is 440,000 Hz; and B4 is 493,883 Hz. While the same notes on other instruments can be played at the same frequencies, it is also understood that some instruments naturally play in one key or another.
The term dissonant note as used in this document is a note that is not within a correct key or chord, where the correct musical key and the correct chord are the musical key or chord that is currently being played by another musician. or musical fountain.
j MPI
The term blue note (blue note) as ^ amiisa ^ l; i KMI<sup>:</sup>kIA<sup>;</sup>. * .— this document is a note that is not within a correct musical key or chord, but is allowed to be played without modification.
h The term musical accompaniment entry note as used herein, is a note played by an accompanying musician that is associated with a note played on the corresponding main melody.
General description of the invention
The following briefly describes various embodiments in order to provide a basic understanding of some aspects of the invention. This brief description is not intended to give a long picture. It is not intended to identify key or critical elements, or to delineate or otherwise narrow the scope. Its purpose is merely to present some concepts in a simplified form as a prelude to the more detailed description that follows.
In short, various modalities are directed toward generating a multi-track recording by looping through a set of previously recorded audio tracks and receiving new audible input for each added audio track. In one embodiment, each of the audio tracks in the recording of
IMPI
INSTITUTO MEXICANC multiple tracks can be generated from dfi ^ wSSÉ ^ a ^ audible vowel of an end user. Each new input aiidl-bl-e can be provided after the current recording is played repeatedly, or by looping, one or more | b times. This recording sequence, separated by looping periods during which no new input tracks are received, can allow a user to listen to the current recording in minute detail, continuously and without the time-related pressure of an additional input required immediately. Looping playback, independent of a loop into which an additional track is input, can also allow other actions to be performed, such as modifying a previous track or changing the parameters of the recording system.
Furthermore, at least one of the audio tracks in the multi-track recording may comprise one or more musical instrument sounds generated based on one or more different sounds provided at the audible input. Various forms of processing can be performed on the received audible input to create the audio track, including aligning and adjusting the timing of the audible input, frequency recognition and adjustment, converting the audible input to a timbre associated with a musical instrument. , adding known sound signals associated with the musical instrument, and the like. Further,
ΙΜρτ each of these processes can be executed
OS THE PROH8OAD inoustiial
<img file="MX345588B_D0012.tif" />
allowing near-instantaneous playback of a generated audio track and allowing other audible inputs to be immediately and subsequently received for processing and overlay as an audio track on one or more tracks previously recorded in a multi-track recording.
In one embodiment, the looped or repeated portion of the multi-track recording may comprise a single measure of music. The length of this measure can be determined by a tempo and a beat associated with the track. In another embodiment, the number of measures, the looping point for playback of a multi-track recording, can be dynamic. That is, the repetition of a first audio track in multi-track recording may occur at a different time than that of a second audio track in a multi-track recording. The setting of this dynamic loop repeat point, for example, can be determined automatically based on the length of an audible input for subsequent tracks.
Various modes are also geared towards automatically producing a single, best shot derived from a collection of shots. In one mode, multiple takes of a performance are recorded during one or more sessions on a multi-track recorder. Each shot is partitioned
IΜ ΡI0¾ automatically in segments. The quality of c '· ^ u / y'í'o / .í.
each of the takes is scored, based on selectable criteria, and a track is automatically constructed from the best quality segments of each take. In an 05 mode, a best segment is defined by the segment having the highest score among a plurality of segment ratings.
Various modalities are further directed to protect a musician from playing a dissonant note. In one embodiment, the notes from an accompanying musical instrument are received as well as from a main musical instrument. The notes of the accompanying musical instrument are then modified based on the key, chord, and / or timing of the principal. In one embodiment, a virtual instrument, where the instrument's input keys are dynamically mapped onto safe notes, can be provided. Thus, if the performer of the virtual instrument is accompanying a melody, the virtual instrument can identify the safe notes that comprise notes that are or for the current chord of the melody being accompanied or are in the musical key of the melody.
Device architecture
<img file="MX345588B_D0013.tif" />
Figure 1 shows a modality of the system (100) that can be implemented in a variety of devices (50), which can be, for illustrative purposes, any multipurpose computer (Figure 1A) portable computing device (Figure IB) and / or a dedicated gaming system (Figure 1C). The system (100) can be implemented as any application installed on the device. Alternatively, the system can be operated within an http (hypertext transfer protocol) browser environment, which can optionally utilize network plug-in technology to expand the browser's functionality for enable the functionality associated with the system (100). Device 50 may include many more or fewer components than those shown in Figure 29. However, it should be understood by those skilled in the art that certain components are not necessary to operate the system 100, while others, such as a processor, microphone, video display, and audio speaker are important, if not necessary to practicing aspects of the present invention.
As shown in Figure 29, the device (50) includes a processor (2902), which may be a CPU, in communication with a main memory (2904) through a
W PT · * · - Jjl ζ 3 ··
WSTnynp MEXICAM bus (2906). As could be understood by a techni¿tíwíg @ i ££ ^ matter having the present description, the di ^ u4oa ^ y ^ J ^ £ L__<sub>:</sub>..
"W.
claims before them, the processor (2902) could also comprise one or more general processors, digital signal processors, other specialized processors and / or ASICs (application specific integrated circuits), alone or in combination with each other. others. The device (50) includes a power supply (2908), one or more network interfaces (2910), an audio interface (2912), a display controller (2914), a user input manager (2916), an illuminator (2918), an input / output interface (2920), an optional haptic interface (2922), and an optional global positioning system (GPS) receiver (2924). Device 50 may include a camera (not shown), allowing video to be acquired and / or associated with a particular multi-track recording. Video from the camera, or from another source, may further also be provided to an online social network and / or online music community. Device 50 may also optionally communicate with a base station (not shown), or directly with another computing device. Another computing device, such as the base station, may include additional audio-related components, such as a professional audio processor, generator,
IMPI ^. <Τ! · ΠΠΌ MLXICANc && amplifier, speaker, XLR connectors and / or füéa ^ SrBLu aeS power supply. . Continuing with Figure 29, the power supply (2908) can contain rechargeable or non-rechargeable batteries or it can be provided by an external power source, such as an AC adapter or a power docking bracket that can also supplement and / or recharge the battery. Network interface (2910) includes circuitry for coupling device (50) to one or more 10 networks, and is constructed for use with one or more communication protocols and technologies including, but not limited to, global system for mobile communication ( GSM, Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), User Datagram Protocol (UDP) Transmission Control Protocol / Internet Protocol (TCP / IP), SMS, General Packet Radio Service (GPRS), WAP, Ultra Band Wide (UWB), IEEE 802.16 Worldwide Interoperability for Microwave Access (WiMax), SIP / RTP, or any of a variety of other wireless communication protocols. Consequently, the network interface (2910) may include as a transceiver,
IMPJ® transmitting and receiving device, or network tar (NIC). _________________
An audio interface (2912) (Figure 29) is enabled to produce and receive audio signals such as the sound 05 of a human voice. For example, as shown more clearly in Figures 1A and IB, the audio interface (2912) can be coupled to a speaker (51) and / or a microphone (52) to allow the output and input of music within the system. (100). The display controller (2914) 10 (Figure 29) is enabled to produce video signals to control various types of displays. For example, the video controller (2914) can control the video monitor screen (75), shown in Figure 1A, which can be liquid crystal, gas plasma, or a light-emitting diode-based screen. (LED, for its acronym in English), or any other type of display that can be used with a computing device. As shown in Figure IB, the display controller (2914) may alternatively control a portable, touch-sensitive screen (80), which would also be enabled to receive input from an object such as a stylus or the finger of a computer. human hand through a user input handler (2916) (see figure 31). The keyboard (55) can comprise any input device (eg keyboard, game controller, spin ball and / or mouse) enabled
----———— to receive input from a user. For example> 'U PWOPIkDAU' 'ü.irriiAi V¿®j3S (55) may include one or more push buttons, numeric dials, and / or keys. The keyboard (55) may also include command buttons that are associated with selecting and sending
Jb images.
Device (50) also comprises input / output interface (2920) for communication with external devices, such as a headset, speaker (51), or other input or output devices. The input / output interface (2920) can use one or more communication technologies, such as USB, infrared, Bluetooth ™, or the like. The optional haptic interface (2922) is enabled to provide tactile feedback to a user of the device (50). For example, in one embodiment, 15 as shown in Figure IB, where the device (50) is a mobile or portable device, the optional haptic interface (2922) can be used to vibrate the device in a particular way such as as, for example, when another user of a computing device is calling.
The optional GPS transceiver (2924) can determine the physical coordinates of the device (100) on the Earth's surface, which typically returns a location as latitude and longitude values. The GPS transceiver 25 (2924) may also use other geolocation mechanisms.
ΪΜΡΙ positioning, including, but not. '·' ·? Υ ~ '·' 4ΙΑ?
triangulation, assisted GPS (AGPS),
<img file="MX345588B_D0014.tif" />
E-OTD, CI, SAI, ETA, BSS or similar, to determine the physical location of the device (50) on the surface of R the Earth. In one embodiment, however, the mobile device can, through other components, provide other information that can be used to determine a physical location of the device, including for example, a MAC address, an IP address, or the like.
As shown in Figure 29, the main memory (2904) includes a RAM (2924) (random access memory), a ROM (2926) (read only memory, for its acronym in English) , and other storage media. Main memory 2904 illustrates an example of computer readable storage media for storing information such as computer readable instructions, data structures, program modules, or other data. The main memory (2904) stores a basic input / output system (BIOS) (2928) to control the low-level operation of the device (50). The main memory also stores an operating system (2930) to control the operation of the device (50). It will be appreciated that this component may include a general purpose operating system such as a version of MAC
OS, WINDOWS, UNIX, LINUX, or an operating system
I My specialized PI such as, for example, softwarenndeikxic / sis ^^^ K
Xbox 360, Wii IOS, Windows Mobile TM, iOS, Android, webOS, QNX, or Symbian® operating systems. The operating system may include, or interact with, a Java virtual machine module that enables control of hardware components and / or operating system operations through Java application programs. The operating system may also include a secure virtual container, generally also referred to as a sandbox, which allows the safe execution of applications, for example, Flash and Unity.
One or more data storage modules (132) may be stored in memory (2904) of device (50). As should be understood by a person skilled in the art having the present description, drawings and claims before them, a portion of the information stored in the data storage modules (132) may also be stored on a disk drive or other storage medium. associated with the device (50). These data storage modules (132) can store multi-track recordings, MIDI files, WAV files, samples of audio data, and a variety of other data and / or data formats or melody input data in any of the formats explained above. The data storage modules (132) can also store information that describes different capacities, which can be sent to other digits. ^^^^ Ji 'NSTITUTO MEXICANO> l LA ΓΕΟΜίΡΑΙ (100)
<img file="MX345588B_D0015.tif" />
for example as part of a header .....<sup>You</sup>dW communication, on request or in response to certain events, or similar. Furthermore, the data storage modules (132) can also be used to store social network information including agendas, friend lists, aliases, user profile information, or the like.
Device (50) can selectively store and execute a number of different applications, including applications for use in accordance with system (100). For example, an application to be used in accordance with the system (100) may include an Audio Converter Module (140), Recording Module of a Live Session with Replay Loop (RSLL) (142) , Multiple Take Automatic Composer Module (MTAC) (144), Harmonizer Module (146), Track Sharing Module (148), Sound Finder Module (150), Genre Comparator Module (152), y Chord comparator module (154). The functions of these applications are described in more detail below.
Applications on the device (50) may also include a messenger (134) and a browser (136). The messenger (132) may be configured to initiate and manage a messaging session using any of a variety of
^ .- ι · ι · ιιιιι ΙΤ <II »messaging communications including but
MEXICAN INSTITUTE V e-mail, Short Message Service ^ = ESS $ ¡3 acronym in English), Instant Message (IM, for its initials in_____
<img file="MX345588B_D0016.tif" />
English), Multimedia Messaging Service (MMS), Interactive Internet Chat (IRC), mIRC, RSS feeds, and the like. For example, in one embodiment, the messenger 243 may be configured as an IM messaging application, such as AOL Instant Messenger, Yahoo! Messenger, .NET Messenger Server, ICQ, or similar. In another embodiment, the messenger (132) can be a client application that is configured to integrate and employ a variety of messaging protocols. In one embodiment, the messenger 132 can interact with the browser (134) for message management. Browser 134 can include virtually any application configured to receive and display graphics, text, multimedia, and the like, employing virtually any network-based language. In one embodiment, the browser application is enabled to use Handheld Markup Language (HDML), Wireless Markup Language (WML), WMLScript, JavaScript, Generalized Markup Language Standard (SMGL), Hypertext Markup Language (HTML), Extensible
<img file="MX345588B_D0017.tif" />
<img file="MX345588B_D0018.tif" />
markup (XML, for short display and send a message, variety of other languages
Python, Java, and external network plug-ins can be used.
Device (50) may also include other applications (138), such as computer-executable instructions which, when executed by a client device (100), transmit, receive, and / or otherwise process messages (eg, SMS. , MMS, IM, email and / or other messages), audio, video and allow telecommunication with another user or client device. Other examples of application programs include calendars, search programs, email clients, IM applications, SMS applications, Voice for Internet Protocol (VoIP) applications, contact managers, task managers. , transcoders, database programs, word processing programs, security applications, spreadsheet programs, games, search programs, and so on. Each of the applications described above can be integrated or alternatively, downloaded and executed on the device (50).
Of course, while the various aíhLicaeckoné ^^
IA íBnMJr.al V · .__ previously described are shown as implemented in device (50), in alternative modes, one or more portions of each of these applications can be 05 implemented in one or more remote devices or servers, where the Inputs and outputs of each portion are passed between the device (50) and one or more of the remote devices or servers in one or more networks. Or, one or more of the applications may be packaged for execution on, or downloaded from, a peripheral device.
Audio converter
The audio converter (140) is configured to receive audio data and convert it to a more meaningful form for use in the system (100). One embodiment of the audio converter (140) is illustrated in Figure 2. In this embodiment, the audio converter (140) can include a variety of subsystems including track recorder (202), a track partitioner (204), quantizer 20 (206), frequency detector (208), frequency converter ( 210), instrument converter (212), gain control (214), harmonic generator (216), special effects editor (218), and manual adjustment control (220). The connections for and the interconnections between the various 25 subsystems of the audio converter (140), are not shown
<img file="MX345588B_D0019.tif" />
<img file="MX345588B_D0020.tif" />
<img file="MX345588B_D0021.tif" />
MEXICAN INSTITUTE to avoid making the present invention unclear ^^^ gi however, these subsystems would be electrically ....... and / or logically connected as would be understood by a technician in the field having the present description, the drawings and claims In front of them.
The track recorder (202) allows a user to record at least one audio track either voice or a musical instrument. In one mode, the user can record the track without any accompaniment. However, the track recorder (202) may also be configured to play audio, either automatically or at the request of a user, comprising a metronome track, a musical accompaniment, an initial pitch against which a user can judge its tuning and time, or even previously recorded audio. Metronome track (click track) refers to a periodic tapping noise (such as periodic tapping noise made by a mechanical metronome) that is intended to assist the user in maintaining a consistent tempo. The track recorder (202) may also allow a user to set the length of time for recording - either as a time limit (eg, a number of minutes and seconds) or a number of bars of music. When used in conjunction with the MTAC module (144), as explained below, the track recorder (202) can also be configured to graphically indicate a score.
ΤΜΡΪ associated with various portions of the grab track $$ &% £ paa ^ to indicate, for example, when a user is out of tune, or the like.
In general, a musical compilation is made up of 05 multiple lyric sections. For example, Figure 3 illustrates a typical progression for a pop song that begins with an intro section, followed by alternating stanza and chorus sections, and a bridge section prior to the final stanza. Of course, although not shown, 10 other structures such as choruses, endings, and the like, can be used as well. Thus, in one embodiment, the track recorder (202) may also be configured to allow the user to select the section of a song for which the recorded audio track will be used. These sections can be arranged in any order (either automatically (based on a determination by the genre comparator module 152 or as selected by the end user) to create a complete musical compilation.
Track partitioner 204 divides a recorded audio track into separate partitions that can potentially be addressed and stored as separate sound clips or files that can be processed individually. The partitions are preferably chosen so that the spliced end-to-end segments result in little or no audio artifact. For example, suppose that a
ΙΜΡΙ audible input comprises the phrase pum pa pum'4vtA. | J ^ D «nl (^ modality, the division of this audible input can identify or distinguish each syllable of this audible input into separate sounds, such as pum, pa and pum. However, it should be understood that this phrase can be delineated in other ways, and a single partition can include more than one syllable or word. Four partitions (numbered 1, 2, 3, and 4) each including more than one syllable are illustrated on screen (75) in Figures 1A, IB and
1 C. As illustrated, partition 1 has a plurality of notes that can reflect the same plurality of syllables that have been recorded by the track recorder (202) using microphone inputs (52) from a human or musical instrument source.
To perform the division of an audible track into separate partitions the partitioner (204) may use one or more processes running on the processor (2902).
In an exemplary embodiment illustrated in Figure 4, the track partitioner (204) may include a silence detector (402), stop detector (404), and / or manual partitioner (406), each of which It can be used to divide an audio track into N time aligned partitions. Track partitioner 204 may use silence detector 302 to divide a track wherever silence is detected for a certain period of time.
<img file="MX345588B_D0022.tif" />
That silence can be defined by a threshold of ^ vcmume such that when the audio volume falls below ~~ the defined threshold for a defined period of time, the location on the track is considered silent. The volume threshold and the time period can both be configurable.
On the other hand, the stop detector 404 can be configured to use speech analysis, such as formant analysis, to identify vowels and consonants in the track. For example, consonants like T, D, P, B, G, K, and 10 the nasal ones are delimited by stops in the flow of air in their vocalization. The location of certain vowels or consonants can then be used to preferentially detect and identify break points. Similar to silence detector (402), the types of vowels and consonants used by stop detector (404) to identify split points can be configurable. Manual partitioner (406) may also be provided to enable a user to manually delimit each partition. For example, a user can simply specify a length of time 20 for each partition causing the audio track to be divided into numerous partitions each of the same length. The user can also be allowed to identify a specific location on the audio track where a partition will be created. Identification can be performed graphically using a pointing device, such as a mouse or game controller, in conjunction with the '^ t ^ p ©<sup>1</sup> from--
<img file="MX345588B_D0023.tif" />
graphical user interface illustrated in figures 1Λ, ID and .........
1 C. Identification can also be performed by pressing a button or key on the user input device, such as the keyboard (55), mouse (54) or game controller (56) during audible playback of the audio track by the track recorder (202).
Of course, although the functions of the silence detector (402), stop detector (304), and manual partitioner (406) have been described individually, it is contemplated that the track partitioner (204) may use any combination of the detector. silence, stop detector and / or manual partitioner to segment or divide an audio track into segments. It would also be understood by one skilled in the art having the present description, drawings, and claims before them that other techniques for partitioning or dividing an audio track into segments may be used.
The quantizer (206) is configured to quantize partitions of a received audio track, which may utilize one or more processes running on the processor (2902). The quantization process, as the term is used herein, refers to the offset time of each previously created partition (and consequently the notes contained within the ίΜ Ρϊ partition), as may be necessary for the INOtrTiilAI purpose.
the sounds inside the partitions with a certain pulse.
• · -. * T - --.—, - <sub>| a</sub>
Preferably, quantizer 206 is configured to align the start of each partition chronologically with a previously determined pulse. For example, a metric can be provided where each measure can comprise four beats and the alignment of a separate sound can occur relative to quarter-beat increments, thus providing sixteen points in each four-beat measure for which a partition can be aligned. Of course, any number of increments for each measure (such as three beats for a waltz or polka effect, two beats for a swing effect, etc.) and beat can be used and, at any time during the process, they can be used. adjusted either manually or automatically by a user based on certain criteria such as a user's selection of a certain style or genre of music (eg, blues, jazz, polka, pop, rock, swing, or waltz).
In one embodiment, each partition can be automatically aligned by quantizer 206 with an available time increment for which it was most closely received in recording time. That is, if a sound starts between two time increments in the pulse, then the sound's playback time will be chronologically shifted forward or backward for either.
<img file="MX345588B_D0024.tif" />
of these increments for which your initial startup time is closest. Alternatively, each sound can be automatically shifted in time for each time increment immediately preceding the relative time 05 in which the sound was initially recorded.
In even another mode, each sound can be automatically shifted in time for each time increment that immediately follows the relative time that the sound was initially recorded. A time offset, if any, for each separate sound may also be alternatively or additionally influenced based on a selected genre for multi-track recording, as explained below with respect to the genre comparator (152). In another embodiment, each sound can also be automatically aligned in time with a previously recorded track in a multi-track recording, allowing for a karaoke-like effect. In addition, the length of a separate sound can be greater than one or more time increments, and the quantizer (206) time offset can be controlled to prevent separate sounds from being offset in time so that they overlap within the same track. audio.
The frequency detector (208) is configured to detect and identify the tones of one or more separate sounds that may be contained within each
MEXICAN INSTITUTE partition, which can use one or more processes running on the processor (2902). In a mode / · a ........- - pitch can be determined by converting each separate sound to a frequency spectrum. Preferably, this can be achieved using the Fast Fourier Transform (FFT) algorithm such as the FFT implementation by iZotope. However, it should be understood that any implementation of the FFT can be used. It is also contemplated that a discrete Fourier transform (DFT) algorithm can also be used to obtain a frequency spectrum.
For illustration, Figure 5 shows an example of a frequency spectrum that can be produced by the output of an FFT process executed on a portion of a received audio track. As can be seen, the frequency spectrum (400) includes a major peak at a single fundamental frequency (F) (502) that corresponds to the tone, in addition to the harmonics that are excited at 2F, 3F, 4F ... nF. Additional harmonics are present in the spectrum because, when an oscillator such as a vocal cord or a violin string is excited at a single pitch, it typically vibrates at multiple frequencies.
In some cases, identifying a tone can be difficult due to additional noise. For example, as shown in Figure 5, the frequency spectrum can
<img file="MX345588B_D0025.tif" />
include noise that occurs as a result of audio being from a real-world oscillator such as a voice or instrument, and appears as low-amplitude peaks spread across the spectrum. In one embodiment, this noise can be extracted by filtering the FFT output below a certain noise threshold. Identifying the pitch can also be complicated in some instances by the presence of vibrato. Vibrato is a deliberate modulation of frequency that can be applied to a performance, and is typically between 5.5 Hz and 7.5 Hz. As with noise, vibrato can be filtered from the FFT output by applying a band-pass filter on the frequency mastery, but filtering out vibrato may not be desirable in many situations.
In addition to the frequency domain approximations discussed above, it is contemplated that the pitch of one or more sounds in a partition could also be determined using one or more time domain approximations. For example, in one embodiment, the pitch can be determined by measuring the distance between zero crossover points of the signal. Algorithms such as AMDF (mean magnitude difference function), ASMDF (mean square mean difference function), and other similar auto-correlation algorithms can also be used. .
For effective industrial tuning judgments, the content of a certain tone can also be grouped into notes (of constant frequency) and glissandos (of frequency that is steadily increasing or decreasing). However - unlike fretted or keyed instruments that naturally produce consistent and discrete tunings - the human voice tends to slip and oscillate between notes continuously, making conversion of discrete tunings difficult. Accordingly, the frequency detector (208) can also preferably use pitch pulse detection to identify shifts or changes in pitch between separate sounds within a partition.
Tuning pulse detection is an approximation of tuning event delineation that focuses on the ballistics of the control loop formed between the singer's voice and their perception of it. Generally, when a singer makes a sound, the singer hears that sound a moment later. If Singer 20 hears that the pitch is incorrect, he will immediately shift his voice toward the intended pitch. This negative feedback loop can be modeled as attenuated harmonic motion driven by periodic pulses. Thus, a human voice can be considered a single oscillator: the vocal chord. An example illustration of a pitch changing
<img file="MX345588B_D0026.tif" />
Figure 6. The tension in the vocal cord controls the pitch, and this pitch change can be modeled by the response to a step function such as the step function (604) in Figure 6. Thus, the start A new tuning event can be determined by finding the beginning of the attenuated harmonic oscillation in the tuning; and observing the successive inflection points of the tuning converging to a constant value.
After the pitch events within an audio track partition have been determined, they can be converted and / or stored into a morphology, which is a graph of pitch events over time. An example of a morphology (without partitions) is shown in Figure 7. The morphology can therefore include information identifying the onset, duration, and pitch of each sound, or any combination or subset of these values. In one embodiment, the morph can be in the form of MIDI data, although a morph can refer to any representation of pitch in time, and is not limited to semitones or any particular measure. For example, other such examples of morphologies that can be used are described in Larry Polansky's Morphological Metrics, Journal of New Music Research, volume 25, pp.
<img file="MX345588B_D0027.tif" />
<img file="MX345588B_D0028.tif" />
289-368, ISSN: 09929-8215, which is present as a reference document.
Frequency shifter (210) may be configured to shift the frequency of an audible input, which may utilize one or more processes running on processor (2902). For example, the frequency of one or more sounds within a partition of an audible input can be automatically raised or lowered in order to align it with the fundamental frequency of audible inputs or separate sounds that have been previously recorded. In one embodiment, the determination to raise or lower the frequency of an audible input depends on the closest fundamental frequency. In other words, assuming the composition was in the key of C major, if the audible frequency captured by the track recorder (202) is 270,000 Hz, the frequency shifter (210) will move the note down to 261,626 Hz ( Center C), whereas if the audible frequency captured by the track recorder (202) is 280,000 Hz, the frequency shifter (210) would move the note up to 293,665 Hz (or the D above the center C). Although the frequency shifter (210) first adjusts the audible input to the nearest fundamental frequency, the shifter (210) can also be programmed to make different decisions in risky situations (for example, where the audible frequency is approximately at half between two notes) musical tonality, genre and / or chord. In one mode, the frequency shifter (210) can adjust audible inputs to other fundamental frequencies to achieve greater musical sense based on key, genre and / or chord based on the controls provided by the genre comparator (260) and / or the chord comparator (270), as explained later. Alternatively or additionally the frequency shifter (210) -in response to the instrument converter rail input (212) may also individually shift one or more portions of one or more partitions to correspond to a predetermined set of frequencies or semitones. such as those typically associated with a selected musical instrument, such as a piano, guitar or other stringed, wood or metal instrument.
The instrument converter (212) may be configured to perform the conversion of one or more portions of the audible input to one to more sounds that have a timbre associated with a musical instrument. For example, one or more sounds in an audible input can be converted to one or more instrument sounds of one or more different types of percussion instruments, including a rattle drum, cowbell, bass drum, triangle, and the like. In one embodiment, converting an input
IMPI &
_ _ __ _ _ ______ □ INSTITUTO MACANO audible to one or more sounds of the corresponding instruments of very sticky detail can comprise adapting the timing and amplitude of one or more sounds in the audible input to a corresponding track comprising one or more sounds of the 05 instrument percussion, the sound of the percussion instrument comprising the same or similar timing and amplitude as one or more sounds of the audible input. For other instruments that can play different notes, such as a trombone and other types of brass, 10-string, woodwind, or similar instruments, the instrument conversion can further correlate one or more frequencies of audible input sounds with one or more sounds. with the same or similar frequencies played by the instrument. Furthermore, each conversion can be derived and / or limited by the physical capacities corresponding to the instrument currently being played. For example, the frequencies of instrument sounds generated for an alto saxophone track can be limited by the current frequency range of a traditional alto saxophone. In one embodiment, the generated audio track may comprise a representation in format
MIDI of the converted audible input. The data for the different instruments used by the instrument converter 212 would preferably be stored in memory 2904 and can be downloaded from an optical or magnetic medium, removable memory, or via the network.
<img file="MX345588B_D0029.tif" />
<img file="MX345588B_D0030.tif" />
The gain control (214) can be coñ? SJguasSíSaíf¡ '' 'NDUmiAJ automatically adjust the relative volume of the audible input based on the volume of another of the previously recorded tracks and can use one or more processes running on the processor (2902 ). The harmonic generator (216) can be configured to incorporate harmonics to an audio track, which can use one or more processes to be executed in the processor (2902). For example, additional frequencies, different from the audible input signal, can be determined and added to the generated audio track. Determination of additional frequencies may also be based on the gender of the gender comparator 260 or through the use of other, predetermined parameter settings entered by a user. For example, if the selected genre was a waltz, additional frequencies can be selected from harmonious major chords for the lead music in the lower octave immediately below the lead, in% time with an oom-pa-pa rhythm as follows :
5 5 5 fundamental 3 3, fundamental 3 3. The special effects editor (218) can be configured to add various effects to the audio track, such as echo, reverb, and the like preferably using one or more processes running on the processor. (2902).
<img file="MX345588B_D0031.tif" />
IMPI
He
MEXICAN INSTITUTE DE LA PROPIEDAD audio converter (140) can also<sup>0</sup> include<sup>-</sup>a manual adjustment control (220) to allow a user to manually alter any of the settings automatically configured by the modules discussed above. For example, the manual trim control (220) may allow a user to alter the frequency of an audio input, or portions thereof; allow a user to alter the start and duration of each separate sound; increase or decrease the gain for an audio track; selecting a different instrument to be applied to the instrument converter (212), among other options. As would be understood by one skilled in the art having the present description, drawings, and claims before them, this manual adjustment control (220) may be designed for use with one or more graphical user interfaces. A particular graphical user interface will be explained later in conjunction with Figures 13 A, 13 B and 13 C.
Figure 8 illustrates one embodiment of a file structure for a partition of an audio track that has been processed by the audio converter 140, or otherwise downloaded, ingested, and obtained from another source. As shown, in this mode, the file includes metadata associated with the file, the morphology data obtained (for example, in MIDI format), and the raw audio (for example, in .wav format). The
<img file="MX345588B_D0032.tif" />
<sup>49</sup> IMPI
MEXICAN INSTITUTE OS LA MWtFOAD INDUSTRIAL metadata may include information indicating a profile associated with the creator or provider of the audio track partition. It can also include additional information regarding the audio signature of the data, such as the key, tempo, and partitions associated with the audio. The metadata can also include information regarding potentially available pitch offsets that can be applied to each note in the partition, the amount of time offset that can be applied to each note, and the like. For example, it is understood that, for live recorded audio, there is a possibility of distortion if a pitch is shifted by more than one semitone. Consequently, in one mode, a constraint must be placed on live audio to prevent shifts greater than one semitone.
Of course, different settings and different restrictions can also be used. In another embodiment, the ranges for potential pitch shift, time shift, etc., may also be altered or set by a creator of a partition of an audio track, or any individual with substantial rights to that partition of the audio track. audio, such as an administrator, a collaboration team and the like.
Live loop replay recording session
<img file="MX345588B_D0033.tif" />
The Session Recording Loop Live (RSLL) module (142) implements a digital audio workstation that, in conjunction with the audio converter (140), enables input recording
<img file="MX345588B_D0034.tif" />
audible, generating separate audio tracks, and creating multi-track recordings. Thus, the module
RSLL (142) can allow any of the recorded audio tracks, whether spoken, sung, or otherwise, to be combined with previously recorded tracks to create a multi-track recording. As explained below, RSLL module 142 is preferably also configured to loop repeat at least one measure of a previously recorded multi-track recording for repeat playback. This repeat playback can be performed while new audible inputs are being recorded or the RSLL module (1412) is otherwise receiving instructions for a recording session that is currently being conducted. As a result of this, the RSLL module (142) allows a user to continue editing and composing music tracks while playing back and listening to previously recorded tracks. As will be understood from the explanation below, continuous looping of previously recorded tracks minimizes the user's perception of any latency that may result from the processes that are applied to an audio track that is currently being recorded by the user, and this is how these processes are preferably completed.
Figure 9 illustrates a logic flow diagram showing generally one embodiment of a general information process for creating a multi-track recording using the RSLL module (142) in conjunction with the audio converter (140). In general, the operations of Figure 9 represent a recording session. Said session can be newly created and completed each time a user 10 uses the system (100), and for example, the RSLL module (142). Alternatively, a previous session can be continued and certain elements thereof, such as a previously recorded multi-track recording or other user-specified recording parameters, can also be loaded and applied.
In any of these arrangements, the process 900 begins, after a start block, a decision block 910, where a user determines whether a currently recorded multi-track recording should be played. The process of playing back the current multitrack recording, while allowing other actions to be performed, is generally referred to herein as live loop replay. The content and duration of a portion of the multi-track recording currently being played, without a
<img file="MX345588B_D0035.tif" />
IMPI
CRSTíTím MÍXÍCANO
DELA PkOHílWIMDUS<sup>1</sup> Explicit repetition, called repetition, in a live loop. During playback, the multi-track recording may be accompanied by a metronome, which generally comprises a separate audio track, not ^ 5 stored with the multi-track recording, that provides a series of equally spaced reference sounds or beats that audibly indicate a speed and a measure for a track for which the system is currently configured to record.
In an initial run of the process 900, an audio track may not have been generated yet. In such a state, the playback of an empty multitrack recording in block 910 can be simulated and the metronome can provide the only reproduced sounds for the user.
However, in one embodiment, a user may select to mute the metronome, as explained below with respect to block (964). Visual cues can be provided to the user during recording along with audio playback. Even though an audio track has not been recorded, and the metronome is silent, simulated playback indication, and the current playback position may be limited exclusively to these visual cues, which may include, for example, a changing display. of a progress bar, pointer, or other graphic indication (see Figures 12A, 12B, and 12C).
<img file="MX345588B_D0036.tif" />
<img file="MX345588B_D0037.tif" />
INSTITUTO M'.XICANU
DULA 'KOHE'JA ·
IN3USTUIAL
Recording of multiple live looped tracks in decision block 910 may comprise one or more audio tracks that have been previously recorded. Multi-track recording can include a total length ^ 5 as well as a length that is played back as a live loop. The length of a live loop can be selected to be less than the overall length of a multi-track recording, allowing a user to separately layer different bars of the multi-track recording. The length of a live loop, relative to the total length of a multi-track recording, can be selected manually by a user or alternatively, determined automatically based on the audible input received. In at least one embodiment, the total length of the multi-track recording and the live loop can be the same. For example, the length of the live loop and multi-track recording can be a single measure of music.
When the multi-track recording is selected for playback in block 910, additional visual cues, such as a visual representation of one or more of the tracks, may be provided in sync with the audio playback of a loop in live comprising at least a portion of the multi-track recording 25 reproduced by the user. While recording μ IMPIOS
MSXiCAÑo INSTITUTE
Df LA froheiad vvr<sup>s</sup>'' js ^ r # industrial ^ ¡jW ^ igSS multiple tracks is played, process (900) continues in decision block (900) continues in decision block (920) where a determination is made by an end user if an audio track for multi-track recording 05 will be generated. Recording can be initiated based on the reception of an audible input, such as an audible vocal input generated by an end user. In one embodiment, a sensed amplitude of an audible input may trigger the sampling and storage of an audible input signal received at the system (100). In an alternate embodiment, the generation of a track can be initialized by a manual input received by the system (100). Additionally, generating a new audio track may require both a detected audible input, such as from a microphone, and a manual indication. If a new audio track is to be generated, processing continues at block 922. If the generation of an audio track is not started, the process (900) continues in the decision block (940).
In block 922, an audible input is received by the track recorder 202 of the audio converter 140 and the audible input is stored in memory 2904 on one or more data storage modules ( 132). As used herein, audible refers to a property of an input to the device (50) where, while the input is being provided, it may be
<img file="MX345588B_D0038.tif" />
<img file="MX345588B_D0039.tif" />
INSTITUTE MEX: CANo Oii LA IWHiOAU IHOUrrklAL concurrently, naturally and directly to be heard by at least one user without amplification or other electronic processing. In one embodiment, the length of the recorded audible input can be determined based on the amount R remaining within a live loop when the audible input is first received. That is, the recording of an audible input can be ended after a length of time at the end of a live loop, regardless of whether a detectable number of audible inputs is still being received. For example, if the length of the loop is one measure of four beats per measure and the reception of the audible input is first detected or triggered at the beginning of the second beat, then three beats worthy of being part of the audible input can be recorded, corresponding to the second, third and fourth beats of the measure and thus, these seconds, third and fourth beats would be looped over in playback of the continuously processed multitrack recording in block 910. In such an arrangement, any audible input received after the end of the measure can only be recorded and processed as a basis for another separate track for multi-track recording. Such additional processing of the separate track can be represented as a separate iteration through blocks (910, 920, and 922).
In at least one alternative embodiment, the length of the looping may be dynamically adjusted based on the length of the audible input received at block 922. That is, the audible input could automatically result in an extension of the track length of the multi-track recording that is currently being played at block 910. For example, if the additional audible input is received after a length of a current live loop has been played, then this longer audible input can also be recorded and held to be derived as the new audio track. . In such an arrangement, previous tracks from the multi-track recording can be repeated within subsequent live loops in order to match the length of the received audible input. In one embodiment, the repeat of the shorter previous multitrack recording can be performed a full number of times. This full number of repeats preserves the relationship, if any, between multiple measures of the shortest 20 multitrack recording previously recorded. In this way, the looping point of a multi-track recording and the live loop can be dynamically altered.
Similarly, the length of the track received at block 922 may be shorter than the length of the loop 25 of the live loop that is currently playing (for example, receiving only one measure from the start. .. ñr <sup>d Ί</sup>'<sup>h 1</sup> é during playback of a four-measure-long live loop). In such an arrangement, the end of the audible input can be detected when no additional audible signal has been received after a predetermined time (e.g., a selected number of seconds) following the reception and recording of an audible input of at least one volume threshold. In one embodiment, the detection of this silence may be based on the absence of input above the volume threshold of the current live loop. Alternatively, or additionally, the end of an audible input may be signaled by the reception of a manual signal. The associated length of this shorter entry can be determined in terms of a number of measures with the same number of beats as multi-track recording. In one mode, this number of measures is selected as a factor of the length of the current live loop. In each case, an audible input, once converted to a track in block (924), can be selected manually or automatically for repeat for a sufficient number of times to match a length of the multi-track recording being played. currently playing.
In block 924, the received audible input can be converted to an audio track by the converter. <sup>58</sup> LMP-W
INSTITUTO .MEXICANO «LA PROPERTY«
INDUSTRIAL (140). As explained above, the audio conversion process can include various operations including divide, quantize, frequency and offset detection, instrument conversion, gain control, harmonic generation, adding special effects, and manual adjustment. The order of each of these conversion operations can be altered, and can, in at least one embodiment, be configured by an end user. Also, each of these operations can be selectively applied, allowing the audible input to be converted to an audio track with as much or as little additional processing as required. For example, instrument conversion may not be selected, thus allowing one or more original sounds from an audible input to be substantially included with their original timbre in the generated audio track. At block 924, an echo cancellation process can be applied to filter the audio from other tracks that are being played during the live loop replay of the audio track that is actively being recorded. In one embodiment, this can be accomplished by identifying the audio signal being played during live looping, determining any delay between the output audio signal and the input audio signal; filtering and delaying the output audio signal to resemble the audio signal of
<img file="MX345588B_D0040.tif" />
<sup>59</sup> IMPI
MEXICAN INSTITUTE LA PEOHECAD INDUSTKÍAt entrance; and subtracting the output audio signal from the input audio signal. A preferred echo cancellation process that can be used is one implemented by iZotope, although other implementations can also be used.
The processes of block 924 may subsequently be applied or removed as will be explained later herein with respect to block 942. After converting the audible input to an audio track generated at block 924, processing 900 continues at block 926.
In block 926, the audio track generated from block 924 can be added in real time to a multi-track recording. This can be multi-track already started or, alternatively, a new one with the audio track included as the first track of it. After block 926, process 900 can start again at block 910, where the multi-track recording can be played with the most recently generated audio track included. While the operations of (922, 924 and 926) are shown as being executed in series in figure 9, these steps can also be executed in parallel for each received audible input, in order to further enable real-time recording and playback. an audible input signal. During each audible input, such parallel processing can
<img file="MX345588B_D0041.tif" />
<img file="MX345588B_D0042.tif" />
IMPI
MEXICAN INSTITUTE os ι.Λ ι · ί <αΗ80Αη INDUSTIIXL be executed, for example for each identified separate sound of the audible input, although alternate modes may include other, differently sized portions of the audible input signal.
At decision block 940, it is determined whether one or more audio tracks in the multi-track recording will be modified. For example, an input that may be received indicating that an end user wishes to modify one or more of the previously recorded audio tracks. In one embodiment, the indication can be received through manual input. As noted above, this modification can also be performed during playback of the currently recorded multi-track recording, allowing immediate appreciation of a current state of the multi-track recording by the end user. In one embodiment, the indication may include one or more tracks of the multi-track recording to which it is desired to apply a setting. These tracks can also include one or more new tracks manually added to the multi-track recording. If the indication of a track modification is received, the process (900) continues at block (942); otherwise, the process (900) continues in the decision block (960).
In block 942, the parameters of one or more of the previously converted tracks are received and set parameters that can be entered by an end user.
6ΐ IΜ ΡI ^^ 5
MEXICAN INSTITUTE DE LA FMOREÜAD kSw «>> LW INOUSTKIAL
The parameters for the modification can include any AWV • RWM · IIIWlll · Ι I III IT of the adjustments that can be made using the processes of the audio converter (140), which can include among other examples, mute or play only the track, remove an entire track, adjust the attack speed of an instrument in a track, adjust the volume level of a track, set a playback tempo of all tracks in the live loop, add or remove separate sounds of selected time increments from a track; adjust the length of a live loop and / or the total length of a multi-track recording. Adjusting the length of the live loop may comprise altering the start and end points of the loop relative to the entirety of the overall multitrack recording and / or it may also comprise adding more bars to the tracks that are being repeated at that time. moment in a live loop, adding and / or appending previously recorded bars from the multitrack recording with at least a subset of the tracks previously associated with these bars, or deleting bars from a multi-track recording. Adding a new track may require various aspects of this new track to be manually entered by the user. Also in block (942), a search can be directed by an additional clue through the use of the module<sup>62</sup> IΜ ΡI
INSTITUTO MSXICANO υ 'ιλ rsOPifDAL' c *.<sup>; 1</sup>’<sup>Γ</sup> · · * .. τ3? Sound finder (150) to facilitate the reuse of previously recorded end-user audio tracks.
In block 944, the adjusted parameters are applied to one or more tracks indicated in decision block 940. The application may include converting the adjusted parameter to a format compatible with one or more of the adjusted tracks. For example, one or more numerical parameters can be set to correspond to one or more values applicable to the MIDI or other protocol format. After block 944, process 900 can start over at decision block 910, where at least a portion of a multi-track recording that corresponds to the live loop can be played with one or more of the included modified audio tracks.
At decision block 960, it is determined whether a recording configuration will be modified. For example, an input may be received indicating whether a user wishes to modify one or more aspects of the recording settings. This indication can also be received through a manual input. The display may include one or more parameter settings of a recording setup that will be adjusted. If the end user wishes to modify the recording configuration process (900), he continues at block (962); otherwise, process (900) continues at decision block (980).
<img file="MX345588B_D0043.tif" />
<img file="MX345588B_D0044.tif" />
INSTITUTO MÜXICANO DE LA PRCPHDAI? INDUSTRIAL
At decision block 962, the recording system can be calibrated. In particular, the recording circuit, comprising at least one audio input source, audio output source, and R audio track processing components, can be calibrated to determine the latency of the system (100) in conjunction with the device. (50), preferably measured in thousandths of a second, between a reproduction of a sound through the audio output source and the reception of an audible input through the audio input source. For example, if a recording circuit comprises a headset and a microphone, the latency can be determined by the RSLL (142) to improve the reception and conversion of an audible input, particularly the determination of a relative synchronization between the pulses of a multi-track recording being played and an audible input received. After calibration at block 962, if any, process 900 continues to block 964.
In block 964, other parameter settings of the recording system can be changed. For example, the metronome playback can be turned on or off. Also, the default settings for new tracks or new multi-track recordings can be modified, such as a default tempo and a default set of conversions for an audible input for block (294)
<img file="MX345588B_D0045.tif" />
<img file="MX345588B_D0046.tif" />
MEXICAN INSTITUTE
DI LA FROTIEDAD INDUSTUIAt can be provided. The measure of a current multi-track recording can also be changed in block (964). Other settings associated with a digital audio workstation can also be provided so that 05 can be modified by an end user as would be understood by one skilled in the art having the present description, drawings and claims before them. After block 964, process 900 can return to decision block 910, where settings to recording system 10 can be applied to subsequent recording and modification of audio tracks for multi-track recording.
In block 980, it is determined whether the recording session will be ended. For example, an entry indicating the end of a session may be received from a manual entry. Alternatively, the device (50) may indicate the end of the session if, for example, the data storage (132) is full. If an end-of-session indication is received, the multi-track recording can be stored and / or transmitted for further operations. For example, a multi-track recording can be saved to data storage 132 for future recall, review, and modifications in a new session or a continuation of the session in which the multi-track recording was initially created. Recording of
<img file="MX345588B_D0047.tif" />
<img file="MX345588B_D0048.tif" />
MEXICAN INSTITUTE uf. Multi-track NutiSéTiiAL can also be transmitted from one device (50) to another device (50) over a network for storage in at least one remote data repository associated with a user account. A streamed multi-track recording can also be shared via a network server with an online music community or shared in a game hosted by a network server.
If the recording session has not ended, process 10 (900) returns again to decision block (910). Such a sequence of events may represent periods in which a user is listening to a live loop while deciding which, if any, which additional tracks will be generated, or what other modifications, if any, will be performed. It will be understood by one of ordinary skill in the art having the present description, drawings, and claims before them that each block in the flowchart illustration in Figure 9 (and otherwise), and the combinations of blocks in the diagram illustration of flow, can be implemented by computer program instructions. These program instructions can be provided to a processor to produce a machine, such that the instructions, which are executed in the processor, create the means to implement the actions specified in the block or blocks of the flow chart. Computer program instructions can be executed by a processor to cause a series of operational steps to be executed by the processor to produce a computer-implemented process such that the instructions, which are executed 05 in the processor to provide the steps for implement the actions specified in the block or blocks of the flow chart. The computer program instructions may also cause at least one of the operational steps shown in the blocks of the flow chart to be executed in parallel. In addition, some of the steps can also be executed spreading them out over more than one processor as might arise in a multi-processor computing system. Additionally, one or more blocks or combinations of blocks in the illustration of the flow diagram may also be executed concurrently with other blocks or combinations of blocks, or even in a different sequence from that illustrated without departing from the focus and spirit of the invention. Accordingly, the blocks in the flowchart illustration support combinations of means for executing the specified actions, combinations of steps for executing the specified actions, and means for program instructions for executing the specified actions. It will also be understood that each block in the flowchart illustration, and 25 combinations of blocks in the diagram illustration
Dk LA FIOFÍiaxO Sr ^ a? Flow systems can be implemented by ba ^ a'ddb eft— special-purpose hardware systems, which perform specified actions or steps, or combinations of special-purpose hardware and computer instructions.
The operation of certain aspects of the invention will now be described with respect to various display screens that may be associated with a user interface implementing an audio converter (140) and an RSLL module (142). The illustrated embodiments are non-limiting, 10 and non-exhaustive examples of user interfaces that may be employed in association with the operations of the system (100). The different display screens can include many more or fewer components than are shown. Furthermore, the arrangement of the components is not limited to that shown on these screens, and other arrangements are also envisaged, including the arrangement of various components at different interfaces. However, the components shown are sufficient to describe an illustrative embodiment for practicing the present invention.
Figures 10, 10A and 10B together illustrate a user interface that implements RSLL (142) and aspects of the audio converter (140) for recording and modifying tracks of a multi-track recording. The general screen of the interface (1000) can be considered a control space. Each control displayed in the interface can be operated based on a manual input from a ustrarib, "such<sup></sup>such as through the use of the mouse (54), touch screen (80), push-button (pressure pad), or a device enabled to 06 respond and to transmit physical control. As shown, the interface 1000 displays various aspects of a recording session and a multi-track recording generated as a part of this session. The menu file (1010) includes options for creating a new multi-track recording or loading a previously recorded multi-track recording, as would be understood by those skilled in the art having the present description, drawings and claims before them.
The tempo control (1012) displays a multitrack recording speed in beats per minute. The tempo control (1012) can be directly modified manually by a user. The measure control (1014) displays a measure number for multi-track recording. The measure control (1014) can be configured to display a current measure number during a live loop, a total number of measures, or alternatively be used to select a certain number of measures from multi-track recording to display. later on interface (1000).
The beat pulse control (1016) displays a beat number for multi-track recording. The beat control (1016) may be configured to display the total number of beats for each measure, or alternatively, a current beat number during playback of the multi-track recording. The time control (1018) displays a time for multi-track recording This time control (1018) can be configured to display a general time for multi-track recording, a length of time for a selected live loop in that moment, an absolute or relative time during a live loop, or be used to jump to a certain absolute time of a multi-track recording. The operations of the interface controls (1000), such as the controls (1012, 1014, 1016, 1018 and 1021-1026), can be changed in the block (964) of figure 9. The controls (1020) They correspond to the configuration of track and recording settings explained below with respect to blocks 942 and 962 of Figure 9.
Add Track Control (1021) allows a user to manually add a track to a multi-track recording. After control selection (1021), a new track is added to the multi-track recording and the interface is updated to include additional controls (1040-1054) for the added track, the operations of which are explained as follows. The Render WAV control) (1022) generates and stores a WAV file of — at least · a portion of a multi-track recording. The portions of the multi-track recording represented ^ 5 in this WAV file, as well as other storage parameters, can be entered later by a user after selecting the Render WAV control (1022). In addition, other audio file formats besides WAV may also be available through a control such as the control (1022).
The metronome control (1023) toggles the playback of the metronome. The Arming Control (1024) toggles on and off the RSLL component (142) of the recording and the ability of a device to record an audible input. The arming control (1024) allows an end user to talk to other users, practice vocal input, and create other audible sounds during a recording session without having those sounds converted to audible input that is later processed by RSLL 20 (142 ).
The circuit parameter control (1025) allows a user to calibrate the recording circuit parameters as explained below with respect to Figure 11. The slider (1026) allows the volume of the playback of the recording multiple tracks be
<img file="MX345588B_D0049.tif" />
IMPI MEXICAN INSTITUTE DE LA PROPERTY INDUSTRIAL controlled. The playback control (1030) allows the playback of a multi-track recording. This playback is directed in coordination with the additional recording parameters displayed and controlled through
<img file="MX345588B_D0050.tif" />
controls (1012-1018). For example, the playback control (1030) may initiate playback of a multi-track recording from the positions indicated by the controls (1014-1018) and at a tempo displayed on the control (1012). As noted above, this control 1030 also allows additional audible input recording to generate another audio track for a multi-track recording. The position control 1032 can also be used to control a current playback position of a multi-track recording. For example, control (1032) may cause playback to start at the absolute start of a multi-track recording or alternatively the start of a current live loop.
The grid (1050) in the user interface (1000) represents the playback and synchronization of separate sounds within one or more tracks of a multi-track recording, where each row represents an individual track and each column represents a time increment. Each row can, for example, include a box for each beat increment in a single measure. Alternatively, each row can include enough boxes to play increments of time for a total duration of a live Loop. Boxes with a first shading or color in the grid (1050), such as box (1052), may represent relative timing where a sound is played during a live loop, while other boxes, such as boxes (1054 ), each indicates a time increment within a track where a separate sound is not played. A track added by means of a hand control (1021) initially includes boxes such as the box (1054). Selecting a frame, such as frame (1052) or frame (1054, can add or remove a sound from the track in the time increment associated with the selected frame. 15 added sounds via manual input to a frame in the grid (1050) it may comprise a predetermined sound for a selected instrument for the track, or alternatively, a copy of at least one quantized sound from an audible input for a track. This manual operation with the grid (1050) allows an audible input to generate one or more sounds for a track, and even add copies of one or more of these sounds at manually selected locations within the track.
A progress bar (1056) visually indicates a time increment from a current playback position <sup>73</sup> IMPI ^
MEXICAN INSTITUTE M LA nOPISDAH tVjWi ινοογγκιαι of a multi-track recording. Each track in the grid (1050) is associated with a set of track controls (1040, 1042, 1044, 1046, and 1048). The track delete control (1040) allows the deletion of a track from a multitrack recording and can be configured to selectively delete a track from one or more measures of a multitrack recording.
The instrument selection control (1042) allows the selection of an instrument for which the sounds of an audible input are converted into the generated audio track. As illustrated in Figure 10A, a plurality of instruments, including percussion or other types of instruments other than percussion, can be manually selected from a drop-down menu.
Alternatively, a predetermined instrument or a predetermined progression of instruments may be selected or predetermined automatically for each given audio track. When no instrument is selected, each sound in a generated audio track can correspond substantially to the sounds of the original audible input, including a chime for the initial audible input. In one embodiment, an instrument may be selected based on RSLL (142) enabled to automatically convert particular sounds in an audible input to associated instrument sounds based on, for example, a classification of
<img file="MX345588B_D0051.tif" />
frequency for each particular sound.
The Mute / Solo control (1044) silences an associated track or all tracks except the track associated with the control (1044). The velocity control (1046) allows adjustment of an initial attack or attack strength of the instrument sounds generated for a converted audio track, which can influence the peak, duration, release and shape of the total amplitude of the each instrument sound generated for the audio track. Such velocity i can be entered manually or alternatively, extracted based on the properties of the audible input sounds from which one or more instrument sounds are generated. The volume control (1048) allows individual control of the playback volume of each track in a multi-track recording.
Figure 11 illustrates one embodiment of an interface (1100) for calibrating a recording circuit. Interface 1100 may represent an example of a pop-up display screen, or the like, that may appear when control 1025 (see Figure 10A) is selected. In one embodiment, interface 1100 comprises a microphone gain control 1110 that allows adjustment of the amplitude of a received audible input. Superior control
MEXICAN INSTITUTE DE LA PROPIEDAD (1120) and the lower control (1130) and the control 'd'é ^ íttedicC ^ - vida (1140) provide additional control and validation to identify a received signal as if it were an audible input for further processing by the system (100). Circuit calibration initiates a predetermined metronome and can direct a user to replicate the metronome on an audible output signal. In an alternate embodiment, the metronome for calibration may be directly received as audible input by audio input devices 10 such as a microphone, without requiring a user to audibly replicate the metronome. Based on relative timing differences between generating sounds at the metronome and receiving sounds at the audible input, a system latency (1160) can be determined. This latency value can later be used by RSLL (142) to improve the quantization of an audible input and the relative timing detected between the playback of a multi-track recording and an audible output || received for subsequent derivation of an additional audio track 20 to be added to a multi-track recording.
Thus, as illustrated, the interfaces (1000 and 1100) present users with a control space that welcomes and is friendly, powerful and consistent even 25 intuitive to learn, which is particularly important.
INSTITUTO MEXICANC DS LA PROPíEnA »for a user who is not a professional musician or not 'WtcT familiar with digital audio authoring tools.
Figures 12A, 12B, and 12C together further illustrate another example display screen that may be used in association with recorded or modified audio tracks in a multi-track recording. In this example, the audio frequency (current and morphological (frequency shift later by frequency shifter 210), partition, quantization and tempo information are graphically provided in order to provide the user with an even more intuitive experience. For example, returning to FIG. 12A, a control space (1200) is provided for a live loop. The control space includes a plurality of partition indicators (1204) that identify each of the partitions (or musical measures) on the track (in the case of Figures 12 AC, measures 1 through 4 are shown). In one embodiment of the graphical user interface illustrated in Figures 12 AC, vertical k lines (1206) illustrate the pulse within each measure, with the number of vertical lines per measure preferably corresponding to the top number of a measure break. For example, if a musical composition is selected to be written in one measure of each measure, it will include three vertical lines to indicate that there are three beats in the measure or in the partition. In the same modality of the
<img file="MX345588B_D0052.tif" />
user interface illustrated in the figures horizontal lines (1208) can also identify the fundamental frequencies associated with an instrument. selected for which the audible input will be converted.
As illustrated below in the mode of the figures
AC, an instrument icon (1210) may also be provided to indicate the selected instrument, such as the guitar selected in Figures 12 AC.
In one embodiment illustrated in Figures 12A-C, a solid line (1212) represents an audio waveform of a track as recorded by an end user, either vocally or using an instrument; while the plurality of horizontal bars (1214) represents the morphology of the notes that have been generated from the audio waveform by the quantizer (206) and the frequency shifter (210) of the audio converter (140) . As illustrated, each note of the generated morphology has been shifted in time to align with the pulses of each partition and shifted in frequency to correspond to one of the fundamental frequencies of the selected instrument.
As illustrated by comparing Figure 12A with Figure 12B and Figure 12C, the playback bar (1216) may also be provided to identify the specific part of the live loop that is currently being played by the track recorder ( 202) according to the process of FIG. 9. The play bar 1216 therefore moves from left to right as the live loop is played. After reaching end 5 of the fourth measure, the playhead returns to the beginning of measure one and loops sequentially again. The end user can provide additional audio input at any point within the live loop by recording additional audio at the appropriate point in the loop. Although not shown in Figures 12AC, each additional recording can be used to provide a new track (or set of notes) to represent within the live loop. Separate tracks can be associated with different instruments by adding additional instrument icons (1210).
Figures 13A, 13B, and 13C together illustrate an example of a process for manually altering a previously generated note via the Interface of Figures 12A-C. As shown in FIG. 13A, an end user 20 can select a specific note (1302) using a pointer (1304). As shown in Figure 13B, the end user can then drag the note vertically to another horizontal line (1208) to alter the pitch of the dragged note. In this example, the note (1302) is shown as 25 that has been moved to a higher fundamental frequency. It is contemplated that the notes could also move at frequencies that are between the fundamental frequencies of the instrument. As shown in Figure 13C, the timing of a note can also be altered by selecting the end of the note's morphological representation and dragging it horizontally. In Figure 13C, the duration of the note (1304) has been lengthened. As also shown in figure 13C, the result of the lengthening of the note (1304), is the automatic shortening of the note (1306) by the quantizer (206) to maintain the beat time and avoid the overlapping of notes that are being played by a single instrument. As would be understood by those skilled in the art having the present description, drawings and claims before them, the same or similar methodology can be used to shorten the duration of the selected note resulting in the automatic lengthening of another adjacent note and also that the duration of A note can be changed from the start of the morphological representation in the same way illustrated with respect to modifying the tail of this representation. It should be similarly understood by one skilled in the art that the same methodology can be used to remove notes from a track or copy notes to insert them into other parts of the track.
Figures 14A, 14B, and 14C further illustrate another example visual display for use with system (100). In this example, the visual screen allows a user to record and modify a multi-track recording associated with percussion instruments. Returning to FIG. 14A, a control space (1400) includes a grid (1402) that represents the reproduction and timing of separate sounds within one or more drum tracks. As seen in the illustration of Figures 12A-C, partitions 1-4, each with four times are represented in the example of Figure 14A-C. For example, in Figure 14A, the first row of the grid (1402) represents the reproduction and synchronization of sounds associated with a first base drum, the second row of the grid (1402) represents the reproduction and synchronization of sounds associated with a snare drum, the third and fourth rows of the grid (1402) represent the reproduction and timing of sounds associated with cymbals, and the fifth row of the grid (1402) represents the reproduction and timing of sounds associated with a floor tom. As would be understood by a person skilled in the art having the present description, drawings, and claims before them, these particular percussion instruments and their order on the grid (1402) is meant only to illustrate the concept and should not be seen as limiting the concept. to this particular example.
<img file="MX345588B_D0053.tif" />
<—Marry ¿«'-¿Μ
Each box in the grid represents the time increments for sounds associated with the related percussion instrument, where an unshaded box indicates that no sounds will be played in that time increment, and a shaded box indicates that a sound (associated with the timbre related percussion instrument) will be played at the increment time. Thus, Figure 14A illustrates an example where no sounds will be played, Figure 14B illustrates an example where the sound of a base drum will be played at the times indicated by the shaded boxes, and Figure 14C illustrates an example where the sounds of a base drum and a symbol will be played at the times indicated by the shaded squares. For each percussion instrument track, a sound associated with the particular percussion instrument can be added to the instrument track in different ways. For example, as shown in Figure 14B or 14C, a play bar 1404 may be provided to visually indicate a time increment of a current play position of a multi-track recording during live loop replay. Thus, in Figure 14B, the play bar indicates that the first beat of the third measure is currently being played. A user can then be allowed to add a sound associated with a particular percussion instrument at a time.<sup>82</sup> ϊΜΡΙ ^ ΐΝητητο Mexican oelaixomídad INDUSTRIAL particular by recording a sound at the tempo in which the playback bar (1404) is on the box associated with a particular time. In one embodiment, the instrument track with which the sound will be associated can be manually identified by the user by selecting or clicking on the appropriate instrument. In this case, the particular nature and tuning of the user-made sound should not be important, although it is contemplated that the volume of the user-made sound may affect the gain of the associated sound generated for the percussion track. Alternatively, the sound made by a user may be indicative of the percussion instrument with which the sound will be associated.For example, a user may vocalize the sounds boom, tsk, or ka to indicate a base drum, symbol, or floor tom, respectively. . In even another mode, the user can be allowed to add or remove sounds from a track simply by selecting or clicking a square in the grid (1402).
Automatic multi-shot composition module
The MTAC module (144) (figure 1A) is configured to operate in conjunction with the audio converter (140), and optionally with the RSLL (142), to allow the automatic production of one, best take derived from a collection. of shots. One mode of the MTAC module (144) is
<img file="MX345588B_D0054.tif" />
illustrated in figure 15. In this mode, the MTAC module (144) includes a partition qualifier (1702) to qualify the partitions of each recorded audio take and a composer (1704) to assemble the single single, best take based on the scores identified by the partition qualifier (1702).
The partition qualifier (1702) can be configured to qualify the partitions based on any of one or more criteria, which may be used by one or more processes running on the processor (2902). For example, a partition can be graded based on the hue of the partition relative to a selected hue for the overall composition. Often times, a player can sing a note out of tune without knowing it. In this way, the notes within a partition can also be scored based on the difference between the key of the note and the appropriate key for the partition.
In many cases, however, a novice end user may be unaware of what key they want to sing in. Consequently, the partition qualifier (1702) can also be configured to automatically identify a hue, which may be referred to as Automatic Hue Detection. With Automatic Hue Detection, the partition qualifier (1702) can determine the closest hue to which the end user
INSTITUTO MFJO · 'ΚΟΗίΟΛΟ recorded his audio performance. The system (50) can “'Té'SÚ'ltat ~ any note that is outside the automatically detected key · and can also automatically adjust those notes to fundamental frequencies that are in the automatically determined key.
An illustrative process to determine the musical key is represented in Figure 16. As shown in the first block, this process rates the entire track comparing it against each of the musical keys (C, C # / Db, D, D # / Eb, E, F, F # / Gb, G, G # / Ab, A, A # / Bb, B) with the weight given to each fundamental frequency within a key. For example, the array of tonality weights for some arbitrary major tonality can look like this [1, -1, 1, -1, 1, 1, -1, 1, -1, 1, -1, 1], which assigns a weight to each of the twelve notes in a scale starting with C and continuing with D, etc. Assigning weights to each note (or interval from the tonic) works for any type of key. Notes out of key have been given negative weight. While the magnitudes of the weights are generally less important, they can be adjusted to the individual taste of the user or based on inputs from the gender comparator module (152). For example, some notes in the key are more definitive for that key, so the magnitude of their weights could be greater. Also, some notes not
<img file="MX345588B_D0055.tif" />
<img file="MX345588B_D0056.tif" />
MEXICAN INSTITUTE
Dt THE INDUSTRIAL i fcOHEDAD are in the tonality are more common than oTra ^^ ptrettett remain negative but have smaller magnitudes. So, it would be possible for a user or system 100 (based on input, for example, from the gender comparator module 152) to develop a more refined arrangement of tonality weights for a larger tonality that could be [1, -1, .5, -.5, .8, .9, -1, 1, -.8, .9 -.2, .5]. Each of the twelve major keys would be associated with an arrangement of weights. As would be understood by a person skilled in the art having the present description, drawings and claims before them, the minor shades (or any other) can be adjusted by selecting the weights for each arrangement that considers the shades within the tonality with reference to any document showing the relative position of notes within a key.
As shown in the third block of figure 16, the relative duration of each note with respect to the duration of the total passage (or partition) is multiplied by the weight of the pitch class of the note in the key that is being analyzed. currently for the loop to determine the score for each note in the passage. At the start of each passage, the marker is reset, then the scores for each note as it is compared against the current key are added to each other until there are no more notes in the passage and the process returns to start. analyze the
MEXICAN INSTITUTE DE LA PXORIIOAD INDUSTRIAL passage regarding the following tonality. The result of the main loop of the process is a single key score for each note reflecting the cumulative of all keys for each of the notes in the passage. In the last block of the process in Figure 16, the key with the highest score would be selected as the Best Key (ie the most appropriate for the passage). As would be understood by a person skilled in the art, different hues could be related or have scores similar enough to be essentially related.
In one modality, a pitch class of the note in a key, represented by the index value in figure 17, can be determined using the formula: index: = (note, tuning - key + 12)% 12, where note. tuning represents a numerical value associated with a specific tuning for an instrument, where the numerical values are preferably assigned in order of increasing tuning. Taking the example of a piano, which has 88 keys, each key can be associated with a number between 1 and 88 inclusive. For example, key 1 can be A0 double pedal A, key 88 can be C8 eighth number eight, and key 40 can be middle C.
It may be desirable to improve the precision of the determination of musical tonality than that obtained
MEXICAN INSTITUTE DE PROPIEDAD INDUSTRIAL with the above methodologies. Where such improved precision is desired, the partition qualifier (1702) (or alternatively the harmonizer (146) (explained later)) can determine whether each of the four most probable tones (determined by the initial tonality determination methodology (described above) has one or more major or minor modes. As would be understood by technicians in the field having the present description before them, it is possible to determine the major or minor modes of any plurality of probable tones to achieve an improvement in the precision of the tonality with the understanding that the greater the number of probable tones analyzed, the processing requirements will be higher.
Determining whether each of the probable keys has one or more major or minor modes can be made by executing interval profiling on the notes input to the partition qualifier (1702) (or the harmonizer (146) by the main music source (2404) in some modalities). As shown in Figure 16A, this interval profiling is performed using a 12x12 matrix in such a way that it reflects each class of potential tone. Initially, the values in this array are set to zero. Then, for each note-by-note transition in the note collection, the average of the two note lengths is added to each pre-existing matrix value stored at the location defined by the first note's pitch pitchClass: pitchClass of the second note. So, for example if the collection of notes were:
<td>Note</td><td>AND</td><td>D</td><td>C</td><td>D</td><td>AND</td><td>AND</td>
<td>Duration</td><td> 1</td><td> 0.5</td><td> 2</td><td> 1</td><td> 0.5</td><td> 1</td>
This would result in the matrix values shown in Figure 16A. This matrix is then used in combination with a major key interval profile and a minor key interval profile - as explained later - to calculate a sum of minor keys and a sum of major keys. Each of the major and minor key interval profiles is a 12 x 12 matrix - containing each potential tuning class like the matrix in Figure 16A - with each index in the matrix having an integer value between -2 and 2 in such a way that it weighs the value of several tones in each key. As would be understood by those skilled in the art, the values in the interval profiles can be adjusted to a different set of integer values to obtain a different tonality profile. A potential set of values for the interval profile for major tonality is shown in Figure 16B, while a potential set of values for the profile
<img file="MX345588B_D0057.tif" />
Intervals for minor key is shown in Figure 16C.
Consequently, the sums of major and minor keys can be calculated, as follows:
1. Initialize the sums of major and minor keys to zero;
2. For each index in the note transition arrangement, multiply the integer value by the value at its corresponding location in the matrix of interval profiles of the minor keys;
3. Add each product of the sum that is being made of minor keys;
Four. For each index in the note transition arrangement, multiply the stored value by its corresponding location in the matrix of major key interval profiles; Y
5. Add the product to the sum of major tones that is being carried out.
After completing the calculations of these product sums for each index in the matrix, the values of the sums of major and minor shades are compared with the scores assigned to the plurality of the most likely shades determined in the initial shade determination process. and a resolution is drawn up on which "IMPI"
MEXICAN INSTITUTE DE LA PHOWECAO VWbJM IW r> VI? T £ LA l ' <sup>Ul</sup>* key / mode combination is best. After completing the product sum calculations for each index in the array, the values of the major and minor sum of tonalities are multiplied by their corresponding index in the array in each of the interval profiles.
Subsequently, the sum of those products constitutes the final assessment of the similarity that the given set of notes is in the mode. Therefore, for the example established in figure 16A, for the major mode of C (see figure 16 B), we would have: (1.25 * 1.15) + (1.5 * .08) + (.75 * .91) + ( .75 * .47) + (.75 * -. 74) = 1.4375 + .12 + .6825 + .3525 + (.555) = 2.0375. Thus, for C major, the example melody would result in a score of 2.0375.
Then, to determine if the value in this mode is smaller, however, it is necessary to shift the profile from smaller intervals to the relative smaller. The reason for this is that the interval profile is set to consider the tonic of the mode (not the tonic of the key signature) to be our first column or first row. We can understand why this is true by looking at the underlying music. Any given armor can be either major or minor. For example, the major mode that is compatible with the C major key signature is the C major mode. The minor mode that is compatible with the key signature of C major is the minor mode of A (natural).
Since the upper-left numerical value in our minor interval represents the transition from C to C considering the minor mode of C, all comparison indices would be shifted by three steps (or, more specifically, three columns to the right, and three rows below), since the tonic / root of a minor key relative to the tonic / root of the major key is three semitones down. After shifting three steps, the upper left numeric value in our interval profile represents the transition from A to A in the minor mode of A. To run the numbers using our example from Figure 16 A (with this matrix shifted): ( 1.25 * .67) + (1.5 * -. 08) + (.75 * .91) + (.75 * .67) + (. 75 * 1.61) = .8375 + (.12) +. 6825 + .5025 + 1.2075 = 3.11. So to compare the results of the two modes, we need to normalize the two interval matrices. To do this, simply add all the values of the matrix together, for each matrix, and divide the sums. We find that the larger matrix has about 1.10 of the cumulative sum, so we multiply our smallest mode value by that amount to normalize the results of the two modes. Thus, the results from our example would be that the example set of notes is most likely in the minor mode of A, because 3.11 * 1.10 = 3.421, which is greater than 2.0375 (the result for the major mode).
<sup>92</sup> 1 iVi PI
INSTITUTO MEXICANí OS I. A rBOFIEDAI ~ fNnilSTUIAI ^^ Zj_
The same process described above would apply to · «« »Bgiii m<sub>> lw</sub> each key provided that the initial matrix of note transitions is relative to the key being considered. So using figure 16A as a reference, if in a different composition of example the tonality considered is F major, the rows and columns of the initial matrix, as well as the rows and columns of the interval profiles represented by figure 16B and 16C, would start with F and end with E, instead of starting with C and ending with B (as shown in Figure 16A).
In another mode where the end user knows what musical key they want to be in, the user can identify that key in which case, the process of figure 16 will be started for only the only key selected by the end user instead of the twelve indicated keys. . In this way, each of the partitions can be judged against the single predetermined hue selected by the user in the manner explained above.
In another embodiment, a partition can also be evaluated by comparing it against a chord constraint. A chord sequence is a musical restriction that can be used when the user wishes to record an accompaniment. Accompaniments can typically be thought of as arpeggios of the notes on the chord track and can also include the chords themselves. This is of course
<img file="MX345588B_D0058.tif" />
INSTin'TOMSXICAK t> F! A »tn»! Rn. .
notes that are outside the chord are allowed to be played, but these * should typically be evaluated on their musical merits' An illustrative process for rating the harmony quality of a partition based on a chord sequence constraint is shown in Figures 17, 17A and 17B. In the process of Figure 17, a selected chord is scored per pass according to how well that selected chord would harmonize with a partition (or measure) of the audio track. The chord rating for each note is the sum of a bonus and a multiplier. In the second process box (1700), the variables are reset to zero for each note in the passage. Subsequently, the pitch ratio of the note is compared to the currently selected chord. If the note is in the selected chord, the multiplier is set to the chordNoteMultiplier value set in the first box of the process (1700). If the note forms a tritone (that is, a musical interval spanning three full tones) of the fundamental (for example, C is the fundamental of a major chord of C), then the multiplier is set to the value of the tritoneMultiplier (which as shown in figure 17 A is negative, thus indicating that the note does not harmonize well with the selected chord). If the note is one or eight semitones above the fundamental (or four semitones above the fundamental in the case of a minor chord), then the multiplier is
INSTITUTO MIXICANf:
ι ι v <At Α · vA ΓΙ '. · CF LA MOHEDAL set to the nonKeyMultiplier value (which is as shown in figure 17A is negative again, thus indicating that the note does not harmonize well with the selected chord). Notes that do not fall into any of the above categories are assigned a zero multiplier, and therefore have no effect on the chord grade. As shown in Figure 17B, the multiplier is scaled by the duration of the fraction of the passage that the current note occupies. Bonuses are added to the chord grade if the note is at the start of the passage, or if the note is the root of the current chord selected by the analysis. The chord grade with respect to the passage is the accumulation of this count for each note. Once the first selected chord is analyzed, the system (50) can analyze other selected chords (one at a time) using the process (1700) again. The chord rating of each pass through the process (1700) can be compared to each other and the highest score will determine the chord that would be selected to accompany the passage as if it were the one that best matches that passage. As would be understood by those skilled in the art having the present description, drawings and claims before them, two or more chords can be found to have the same score with respect to a selected passage in which case the system (50) could decide between those chords. based on various selections, including, but not limited to, music track genre. It should also be understood by one skilled in the art having the present description, drawings and claims before them, that the qualification set forth above is to some extent an aspect of design selection rather than of the prevailing genre in Western music. Consequently, it is contemplated that the selection criteria for multipliers could be altered for different genres of music and / or the multiplier values assigned to the different multiplier selection criteria in Figure 17 could be changed to reflect different musical tastes without depart from the spirit of the present invention.
In another embodiment, the partition qualifier (1702) can also judge a partition against the collection of certain allowed pitch values, such as semitones as are typical in Western music. However, the quarter tones of other musical traditions (such as those of Middle Eastern cultures) are viewed in a similar way.
In another embodiment, a partition can also be rated based on the quality of the transitions between the different tones within the partition. For example, as explained above, changes in pitch can be identified using pitch pulse detection.
<img file="MX345588B_D0059.tif" />
In one mode, the same pitch pulse detection --- f—<sub>t</sub> j ihm—-<sub>r</sub> ____ can be used to identify the quality of pitch transitions in a partition. In an approximation, the system can use the generally understood concept that attenuated harmonic oscillators generally satisfy the following equation:
cPx dx, _ + 2ζω<sub>0</sub>- 4- ω £ χ = 0 di<sup>z</sup> at where ωθ is the unattenuated angular frequency of the oscillator and ξ is a system-dependent constant called the attenuation coefficient (for a mass on a spring having a spring constant k and an attenuation coefficient c,.) It is understood that the value ζ - c / 2mo<sub>0</sub>.) of the attenuation coefficient ξ critically determines the behavior of the attenuated system (eg, over attenuated, critically attenuated (ξ = 1), or under attenuated). In a critically attenuated system, the system returns to equilibrium as quickly as possible without oscillating. A professional singer, in general, is able to change his pitch with a response that is critically attenuated. By using pitch boost analysis, the true start of the pitch shift event and the quality of the
I Μ Ρ I institute mixio.no de la prowsüad • NlaUSTRiAL pitch change both can be determined. In particular, In particular, the pitch change event is the deduced step function, where the quality of the pitch change is determined by the value of ξ. For example, Figure 19 shows a stepped response of an attenuated harmonic oscillator for three values ξ. In general, values of ξ> 1 denote poor vocal control, where the singer hunts to find the target pitch. Thus, the larger the ξ value, the poorer the grade of the pitch transition attributed to the partition.
Another example method of rating the quality of the pitch transition is shown in Figure 20. In this mode, scoring a partition may comprise receiving an audio input (process 2002), converting the audio input into a morphology of pitch events showing the true oscillations between pitch changes (process 2004), using the morphology tuning events to construct a waveform with critically attenuated pitch changes between each tuning event (process 2006), compute the difference between the pitch in the constructed waveform with the original audio waveform (2008 process), and compute a rating based on this difference (2010 process). In one embodiment, the rating may be based on the square root of the 'mean square error between the filtered pitch ^ ~' 'and' reconstructed. In simple terms, this calculation can tell the end user how far they deviated from the ideal pitch, which in turn can be converted into a pitch transition score.
The rating methods described above can be used to rate a partition by comparing against either an explicit reference or an implicit reference. An explicit reference can be a previously recorded melody track, musical key, chord sequence, or note range. The explicit case is typically used when the player is recording in unison with another track. In the explicit case an analogy could be made to judge how in karaoke the musical reference exists and the track is being analyzed using the melody previously known as the reference. An implicit reference, on the other hand, may be a target melody (i.e., the system's best forecast on the notes the performer intends to produce) computed from multiple previously recorded takes that have been saved by the track recorder (202 ) in the data warehouse (132). The implicit case is typically used when the user is recording the main melody of a song during which no reference is available, such as a<sup>99</sup> IMPI ^
MEXICAN INSTITUTE ns LA rropisoAi O
INOUSTKIA The original composition or song for which the partition qualifier (1702) is not aware.
In the case where a reference is implicit, a reference can be computed from the takes. This is typically achieved by determining the centroid of the morphologies for each of the N partitions of each previously recorded track. In one embodiment, the centroid of a set of morphologies is simply a new morphology constructed by taking the average mean of the pitch and duration for each event in the morph. This is repeated for η = 1 to N. The resulting centroid would then be treated as the morphology of the implicit reference track, an illustration of a centroid determined in this way for a single note is shown in Figure 18, with the dotted line representing the resulting centroid. It is contemplated that other methods can be used to compute the centroid. For example, the modal average value of the set of morphologies for each of the shots could be used instead of the average of the average. In either approach, any isolated values can be discarded before computing the mean or mean. Those skilled in the art having the present description, drawings and claims before them, would understand that additional options can be developed to determine the centroid of the taps based on the
100
<img file="MX345588B_D0060.tif" />
principles set forth above in ion having to direct excessive experimentation.
As would be understood by one skilled in the art having the present disclosure, drawings and claims before them, any number of the foregoing independent methodologies for qualifying partitions can be combined to provide an analysis of a broader set of considerations. Each grade can be given with the same or different weight. If the ratings are given with different weights it may be based on the particular genre of the composition as determined by the genre comparator module (152). For example, in some musical genres a higher value may be placed on one aspect of a performance over another. The selection of which rating methodologies are applied can also be determined automatically or manually selected by a user.
As illustrated in Figure 23, the partitions of a musical performance can be selected from any of a plurality of recorded tracks. The composer (1704) is configured to combine partitions from a plurality of the recorded tracks to create an ideal track. The selection could be manual through a graphical user interface where the user could see the ratings identified by each version of a partition, the audition of each version of a partition, and
<img file="MX345588B_D0061.tif" />
pick a version as the 'best' track. Alternatively, or additionally, the combination of partitions can be performed automatically by selecting the version of each partition of a track with the highest scores based on the concepts entered above.
Figure 21 illustrates an example embodiment of a process for providing a single best take from a collection of takes using the MTAC module (144) in conjunction with the audio converter (140). In step 2102, the user establishes a configuration. For example, the user can select whether a partition will be qualified against an explicit or implicit reference. The user can also select one or more criteria (ie key, melody, chord, target, etc.) to use to rate a partition, and / or provide ratings to identify the relevant weight or importance of each criterion. A take is then recorded in step 2104, partitioned in step 2106, and converted to a morphology in step 2108 using the process described above. If the RSLL module (142) is being used then, as described above, at the end of the take, the track can automatically return to the beginning, allowing the user to record another take. Also, during recording the user can choose to listen to a metronome, a previously recorded track, a MIDI version of any track
102 alone, or a MIDI version of a target track cdffi ^ íitaSt ^ as previously explained with respect to UKa 'r ef ere ne explicit or implicit (see figures 18, 19, 20 and 21). This allows the user to listen to a reference against which he can produce the next (hopefully improved) take.
In one embodiment, the end user can select the reference and / or one or more methods against which the recorded takes should be rated, step (2110). For example, the user settings may indicate that the partition should be scored against a key, a melody, chords, a meta morphology constructed from the centroid of one or more tracks, or any other method explained above. Guide selection can be done manually by the user or set automatically by the system.
The partitions of a take are rated in step 2112, and in step 2114, an indication of the rating for each partition in a track may be indicated to the user. This can benefit the end user by providing an indication of where the end user tuning or timing failed so that the end user can improve on future takes. An illustration of a graphical screen to display the rating of a partition is shown in Figure 22. In particular, in the
<img file="MX345588B_D0062.tif" />
<sup>103</sup> IMPI instituto M £ xican <. . m LA ΜΟΑΙΗΡλι
INDUSTRIAL figure 22 the vertical bars represent an audio waveform as recorded from an audio source, the solid ”line, primarily horizontal, shows the ideal waveform that the audio source was trying to imitate, and the arrows represent as the tuning of the audio source (for example, a singer) varied from the ideal waveform (called explicit reference).
In step 2116, the end user manually determines whether or not to record another take. If the user desires another shot, the process returns to step 2104. Once the end user has recorded all of the multiple takes for a track, the process proceeds to step (2118).
In step 2118, the user may be provided with an option to decide whether an overall track considered the best will be compiled from all the takes manually or automatically. If the user selects to create a manual composition, the user can, in step 2120, simply listen to the first partition of the first take, followed by the first partition of the second take, until each of the first candidate partitions has been heard. An interface that can be used to facilitate the audition and selection between the different takes of the partitions is shown in figure 23 where the end user by means of a pointing device (like a mouse) to click on each track taken for each partition to give
104
WICKED
MEXICAN INSTITUTE places each track to play.And later the user selects one of these candidate partitions as the best execution of this partition by, for example, double clicking on the desired track and / or clicking and dragging the desired track to the bottom of final compiled track (2310). The user repeats this process for the second, third and subsequent partitions, until the end of the track is reached. The system builds a better track by joining the selected partitions into a single, new track in step (2124). The user can then also decide whether to record additional takes in order to improve his performance in step (2126). If the user chooses to compile the best track automatically, a new track is joined in step 2122 based on the scores for each partition in each take (preferably using the highest rated take for each partition).
An example of a virtual best track that is put together from current recorded track partitions is also illustrated in Figure 223. In this example, the final compiled track (2310) includes a first partition (2302) from take 1, a second partition (2304) from track 5, a third partition (2306) from take 3 and a fourth partition (2308) taken from track 2, with no partitions used from track 4.
I -
105
<img file="MX345588B_D0063.tif" />
Harmonizer
The harmonizer module (146) implements a process to harmonize notes from an accompaniment source with a key and / or chord from a main source, which can be a vocal input, a musical instrument (real or virtual), or a melody. pre-recorded that can be selected by a user. An example modality of this process of harmonization with an accompaniment source is described in conjunction with figures 24 and 25. Each of these figures are illustrated as a data flow diagram (DFD). These diagrams provide a graphical representation of the flow of data through an information system, where data items flow from an external data source or internal data repository to an internal data repository or external data collector. through an internal process. These diagrams are not intended to provide information about the timing and ordering of the processes, or whether the processes will operate in sequence or in parallel. Also, control signals and processes that convert input control flows to output control flows are generally indicated by dotted lines.
Figure 24 shows that the harmonizer module (146) may generally include a
106
IMPIg MEXICAN INSTITUTE note (2402), a main music source (^^^ i ^ Eutuní ^ accompaniment source (2406), a C chord / key (2408), and a controller (2410). As shown, the Note transformation can receive main music input from the main music source 2406. The main and backing music can each be made up of live audio or previously stored audio. In one embodiment the harmonizer module (146) may also be configured to generate the backing music input based on a melody from the main music input.
The note transform module (2402) may also receive a musical key and / or a selected chord from the chord / key selector (2408). The control signal from the controller (2410) tells the note transformation module (2402) whether the music output should be based on the main music input, music accompaniment input and / or the music key or chord of the music selector. chord / tonality (2408) and how the transformation should be handled. For example, as described above, the musical key and chords can be derived from either the main melody or the accompaniment source or even from the manually selected key or chord indicated by the chord / key selector (2408).
<img file="MX345588B_D0064.tif" />
MEXICAN INSTITUTE
OF THE RKOPISDAtt INDUSTKIAl
<img file="MX345588B_D0065.tif" />
<img file="MX345588B_D0066.tif" />
Based on the control signal, the txans-f note rmation module (2402) can alternatively transform the main music input to a note consonant with the chord or musical key, producing a harmonious output note. In one embodiment, the input notes are mapped to harmonious notes using a pre-set consonance metric. In an embodiment explained in more detail below, the control signal may also be configured to indicate whether one or more blue notes can be allowed into the backing music input without transformation by the note transformation module (2402).
Figure 25 illustrates a data flow diagram showing generally more detail on the processes that can be executed by the note transformation module (2402) of Figure 24 in selecting notes to harmonize with the main music source ( 2404). As shown, the main music input is received at process 2502, where a note of the main melody is determined. In one embodiment, a note of the main melody can be determined using one of the techniques described, such as converting the main music input into a morphology that identifies its start, duration and pitch or any subset or combination thereof. Of course, as would be understood by a person skilled in the art having regard to the present description, drawings and c
MLXICANO INSTITUTO DE LA PROPERTY INDUSTRIAL - «iu22í— claims before them, other methods of determining a note from the main melody may be used.
For example, if the main music input is already in MIDI format, determining a note may simply include extracting a note from a MIDI stream. As the main melody notes are determined, they are stored in a main music buffer (2510). The proposed musical accompaniment input is received in process 2504 from the accompaniment source 2406 (as shown in FIG. 24). The process (2504) determines an accompaniment note and can extract the MIDI note from the MIDI stream (where available), converts the musical input into a morphology that identifies its start, duration and pitch, or any subset or combination thereof or use another methodology that would be understood by a person skilled in the art having the present description, drawings and claims before them.
In the process (2506), a chord of the main melody can be determined from the notes found in the main music buffer (2516). The chord of the main .melody can be determined by analyzing the notes in the same manner below in association with figure 17 above or using another methodology understood by those of ordinary skill in the art (such as the analysis of a progression using a Chord Pattern<sup>109</sup> IMPI
1NSTTUTG MEXICAN 'CE THE PROPERTY l * JL'U5TSUAL hidden Markov chain as executed by the Chord Comparator (154) described below). The Hidden Markov Chain Model can determine the most likely sequence of chords based on the fe chord harmonization algorithm explained in this document in association with a transition matrix of chord probabilities which is based on diatonic harmony theory. . In this approximation, the probability of a given chord correctly harmonizing one measure of the melody is multiplied by the probability of the transition from a previous chord to the current chord, and then the best trajectory is found. The timing of the notes as well as the notes themselves can be analyzed (among other potential considerations, such as gender) to determine the current chord of the main melody. Once the chord has been determined its notes are passed to the note transform module (2510) to await potential selection by the consonance control control signal (2514).
In process 2508 of FIG. 25, the musical key 20 of the main melody can be determined. In one embodiment, the process described with reference to Figure 16 above can be used to determine the key of the lead tune. In other embodiments, statistical techniques including use of the hidden Markov Chain Model 25 or the like can be used to determine a
<img file="MX345588B_D0067.tif" />
musical tonality from the notes aimanAnadaA an ai, main music buffer. As would be understood by one skilled in the art having the present description, drawings, and claims before them, other methods for determining musical tonality are similarly contemplated, including but not limited to combinations of the process (1600) and the use of statistical techniques. The process output (2508) is one of many inputs to the note transform module (2510).
Process (2510) (see figure 25) transforms the note used as accompaniment. The transformation of the musical accompaniment note input in process 2510 is determined by the consonance control output 2514 (explained in some detail below). Based on the output of the consonance control (2514), the note transformation process (2510) can select between (a) the note input of the process (2504) (which is shown in Figure 24 as having received the input accompaniment source music (2406)); (b) one or more notes of the chord (which is shown in Figure 24 as having been received from the chord / tonality selector (2408)); (c) a note of the selected musical key (the identity of the key having been received from the chord / key selector 2408 (as shown in FIG. 24)); (d) one or more entry notes
111
<img file="MX345588B_D0068.tif" />
MEXICAN INSTITOTE
Say INDUSTRIAL PROPERTY
<img file="MX345588B_D0069.tif" />
process chord (2506) (which is displayed as based on the notes and musical key determined from the notes in the main music buffer (2516)); or € the musical key determined from the notes in the main music buffer (2516) by the process (2508).
In process 2512, the transformed note can be rendered by modifying the note of the musical accompaniment input and modifying the timing of the note of the musical accompaniment input. In one mode, the rendered note is played audibly.
Additionally or alternatively, the transformed note can also be rendered visually.
The consonance control (2514) represents a collection of decisions that the process makes based on 15 or more entries from one or more sources that control the selection of notes made by the note transformation process (2510). The consonance control 2514 receives a number of input control signals from the controller 2410 (see Figure 24), which may come directly from user input (probably from a graphical user input interface or from a preset configuration), the harmonizer module (146), the genre comparator module (152), or another external process. Among the potential user inputs that can be considered by the consonance control (2514) are
112
<img file="MX345588B_D0070.tif" />
find user inputs that require the output note to be (a) restricted to the chord selected by the chord / key selector (2408) (see figure 24); (b) restricted to the key selected by the chord / key h selector (2408) (see figure 24); (c) in harmony with the chord or key selected by (2408) (see figure 24); (d) restricted to the chord determined by process (2506); (e) restricted to the tonality determined by the process (2508); (f) in harmony with the chord or key determined from the main notes; (g) restricted within a certain range of tones (eg, below the middle C, within two octaves of the middle C, and so on); and / or (h) restricted within a certain selection of tones (ie, minor, augmented and so on).
To an approximation, the consonance control 2514 may further include logic to find bad sounding notes (based on the selected chord progression) and move them to the closest pitch of the chord. A bad sounding note might still be in the correct key, but it would sound bad on the chord being played. The notes are categorized into three different sets, related to the chord on which they are played. Sets are defined as chordTones, nonChordTones, and badTones. All the notes would still be in the correct key, but would have varying degrees of how bad
113
<img file="MX345588B_D0071.tif" />
they sound on the chord being played; chordTones sound. better, nonChordTones sound reasonably good, and badTones sound bad. Additionally, a strictness variable can be defined where the notes are categorized based on how strictly they must adhere to the chords. These strictness levels can include: StrictnessLow, StrictnessMedium, and StrictnessHigh. For each level of strictness, the three sets of chordTones, nonChordTones, and badTones vary. Furthermore, for each level of strictness, the 10 three sets are always related to each other in this way; chordTones are always the tones that the chord consists of, badTones are the tones that will sound bad at this strictness level, and nonChordTones are the diatonic tones that remain and have not been counted for any ensemble.
Because chords vary, badTones can be categorized specifically for each level of stringency, while two other sets can be categorized when a specific chord is given. In one embodiment, the rules for identifying bad-sounding notes are static, as follows:
StrictnessLow (badTones):
One on a major chord (for example F over C major);
A 4 ^ augmented over a major chord (for example 25 F # over C major);
114
IΜ ΡI
MEXICAN INSTITUTE
DE LA MOHEDA »s ^ * k¡? £ -ÍS
INDUSTRIAL
A 63 minor over a minor chord (for example, G # over C minor)
A 6 ^ major over a minor chord (for example, A over C minor); Y
A minor 2nd over any chord (for example, C # over C minor or C major).
StrictnessMedium (badTones):
A 4 »over a major chord (for example F over C 10 major);
A 43 augmented over a major chord (for example F # over C major);
A minor 6th over a minor chord (for example, G # over C minor);
A 63 major over a minor chord (for example, A over C minor);
A 23 minor over any chord (for example, C # over C minor or C major); Y
A 73 major over a major chord (for example, B 20 over C).
StrictnessHigh (badTones):
Any note that does not fall into the chord (not a chordTone).
Being a bad note alone cannot be the only basis for correction, the logic of basic counterpoint
115
<img file="MX345588B_D0072.tif" />
<img file="MX345588B_D0073.tif" />
Based on a classical melodic theory you can aat-tt-sado to identify those notes that would sound bad in context. The rules about whether notes are moved to a chordTone can also be defined dynamically in terms of the strictness levels described above. Each level can use the note set definitions described in its corresponding level of stringency, and can also be determined in terms of stepTones. A stepTone is defined as any note that falls directly before a chordTone in time, and is two or less semitones away from the chordTone; and any note that falls directly after a chordTone in beat, and is also two or less semitones away from the chordTone. In addition, each level can apply the following specific rules:
SrictnessLow: For StrictnessLow, stepTones are extended to two notes outside of a chordTone, so that any note that moves to or from another note that moves to or from a chordTone is also considered a stepTone. Also, any note that is a badTone as defined by StrictnessLow is moved to a chordTone (the closest chordTone will always be a maximum of two semitones away in a diatonic setting), unless the note is a stepTone.
StrictnessMedium: For StrictnessMedium, the stepTones do not extend to notes that are two notes away from the chordTones in time, as they are in StrictnessLow.
<img file="MX345588B_D0074.tif" />
<img file="MX345588B_D0075.tif" />
INSTITUTO MÍXICANI DE LA PaONtDAD IMCUSTBIAL
Any note that is a badTone as defined by
StrictnessMedium is moved to a chordTone. Also, any nonChordTone that falls on the first beat of a downbeat is also moved to one note of the chord. The first beat is defined as any note that begins before the second half of any beat, or any note that lasts longer than the entire first half of any beat. The downbeat can be defined as follows:
• For metrics that have a number of beats that is 10 divisible exactly by three (3/4, 6/8, 9/4), every third beat after the first beat, as well as the first beat, is a strong beat (in 9/4, 1, 4 and 7) - • For metrics that are not exactly divisible by three, and are divisible by exactly two, the downbeat is the first beat, as well as every second beat after that (in 4 / 4: 1 and 3; in 10/4: 1, 3, 5, 7, 9).
• For metrics that are not exactly divisible by two or three, and that also do not have five beats 20 (five is a special case), the first beat, as well as every second beat after EXCEPT the second to the last beat is considered as a downbeat (in 7/4: 1, 3, 5).
• If the metric has five beats per measure, 25 the downbeats are considered to be 1 and 4).
117
<img file="MX345588B_D0076.tif" />
<img file="MX345588B_D0077.tif" />
StrictnessHigh: any note that is defined as a badTone by StrictnessHigh is moved to a chordTone. However, if a note is moved to a chordTone, it will not be moved to the third of the chord. For example, if D is moved over the C chord, the note can be moved to C (the root) instead of E (the third).
Another input to the consonance control (2514) is the consonance metric, which is essentially a feedback path for the note transformation process (2510). First of all, a consonance is generally defined as sounds that make a pleasant harmony with respect to a base sound. The consonance can also be thought of as the opposite of dissonance (which includes any sounds used freely even if they are not harmonious). Therefore, if an end user caused the consonance control (2510) to be fed with the control signals through the controller (2410) that restricted the output note of the note transformation process (2510) to the chord or key manually selected using the chord / key selector (2408), then one or more of the output notes may not be harmonious with the main music buffer (2516). An indication that the exit note was
IMPIí ^ a
118 <sup>INSTI</sup>JyTOMMICANC.
»« THE MOTORCYCLE »A · <noustria<sub>(</sub> non-harmonious (i.e., the metric of be eventually fed back to the consonance control (2514). Although, the consonance control (2514) is designed to force the output note track generated by the note transformation (2510) back Due to the latencies inherent in feedback and programming systems, it is expected that a number of non-harmonious notes will be allowed to enter the musical output. In reality, by allowing at least some non-harmonious notes and even non-harmonious breaks in the music produced by the system, it would make it easier for the system (50) to produce a less mechanical sounding form of musical composition, something desired by the inventors.
In one embodiment, another control signal that can also be input to consonance control 2514 indicates whether one or more blue notes can be allowed in the music output. As noted above, the term blue note for purposes of this description has been given a broader meaning than its original use in blues music as a note that is not in a correct key or chord, but is allowed to be touch without transformation. In addition to taking advantage of system latencies to provide minimal insertion of blue notes, one or more blues accumulators (preferably software encoded rather than hardwired) can be used to provide some additional headroom for blue notes. So, for example, one accumulator can be used to limit the number of blue notes within a single partition, another accumulator can be used to limit the number of blue notes in adjacent partitions, yet another accumulator can be used to limit the number of notes. blue for some predetermined time interval or a total number of notes. In other words, the consonance control via consonance metric can be by counting one or more of any of the following parameters: elapsed time, the number of blue notes in the music output, the total number of notes in the music output music, the number of blue notes per partition, and so on. Default limits, automatically determined and determined / adjusted in real time can be programmed in real time or as presets / defaults. These values can also be affected by the genre of the current composition.
In one embodiment, system 100 can also include a super keyboard to provide a source of backing music. The super keyboard can be a physical hardware device, or a graphical representation that is generated and displayed by a computing device. In either mode, the super keyboard can be designed as manual input for the chord / tonality selector.
120
MEXICAN INSTITUTE DE LA PtOPUDAn> | N »U $ TRIAL (2408) of Figure 24. The super layout preferably includes at least one row of input keys on a keyboard that dynamically maps to notes that are in a musical key and / or that they are in a chord (that is, they are part of the chord) with respect to the existing melody. A super keyboard can also include a row of input keys that are not harmonious with the existing melody. However, the non-harmonious input keys pressed on the super keyboard can then be dynamically mapped to the notes that are in the musical key of the existing melody, or to notes that are chord notes for the existing melody.
One embodiment of a super keyboard in accordance with the present invention is illustrated in Figure 26. The embodiment 15 illustrated in Figure 26 is shown with respect to the notes for a standard piano, although it could be understood that the super keyboard can be used by any instrument. In the embodiment shown in Figure 26, the top row 2602 of input keys of a super keyboard maps onto 20 the notes of a standard piano; the middle row (2604) maps over the notes that are in a musical key to the existing melody; and the bottom row (2606) maps onto notes that are within the current chord. More specifically, the top row exposes twelve notes by 25 octaves as on a regular piano, the middle row exposes ι · * τ ι jl-l 1-1 i> ni iai ri ii mi ι i · n <sup>121</sup>
INSTITUTO MSXICAH · OS U nOPfEOAO O »*« iaí $ U
INDUSTRIAL eight notes per octave and the row below displays three notes per octave. In one embodiment, the color of each input key in the middle row may depend on the current musical key of the melody. As such, when the musical key of the melody changes, the input keys that were selected to be displayed in the middle row also change. In one embodiment, if a non-harmonious musical note is entered by the user in the top row, the super keyboard can also be configured to automatically play a harmonious note instead. In this way, the performer can accompany the main music in an increasingly restricted way the lower the row he chooses. However, other provisions are also foreseen.
Figure 2-7A illustrates one embodiment of a chord selector in accordance with the present invention. In this mode, the chord selector may comprise a graphical interface of a circle of fifths (2700). The circle of fifths (2700) shows the chords that are in the 20 musical key with respect to the existing melody. In one mode, the circle of fifths (2700) displays chords derived from the currently selected musical key.
In one embodiment, the currently selected musical key is determined by the melody, as explained above. Additionally or alternatively, the circle
<img file="MX345588B_D0078.tif" />
outermost concentric of the circle of fifths provides a. mechanism to select a musical key. In one mode, a user can enter a chord using the chord / key selector (2408), selecting a chord from the circle of fifths (2700).
In one mode, the circle of fifths (2700) shows seven chords related to the currently selected musical key — three major chords, three minor chords, and one diminished chord. In this mode, the diminished chord is located in the center of the circle of fifths; the three minor chords surround the diminished chord; and the three major chords surround the three minor chords. In one mode, a performer is allowed to select a musical key using the outermost concentric circle, where each of the seven chords represented by the circle of fifths is determined by the selected musical key.
Figure 27B illustrates another potential embodiment of a chord selector in accordance with the present invention at a particular time during operation of the system (50). In this embodiment, the chord selector may comprise a chord flower (2750). Similar to the circle of fifths (2700), the chord flower (2750) displays at least a subset of the chords that fall musically within the current musical key of the current audio track. And the flower of
<img file="MX345588B_D0079.tif" />
<img file="MX345588B_D0080.tif" />
chords (2750) also indicates the chord that is currently being played. In the example illustrated in figure 27B, the key is C major (as can be determined from the identity of the major and minor chords included in the flower petals and in the center) and the chord played at that time it is indicated by the chord shown in the center, which at the illustrated playback time is C major. The chord flower (2750) is arranged to provide visual cues as well as the probability of any displayed chord immediately following the chord played at that time. As illustrated in Figure 27B, the most likely chord progression would be from currently played C major to G major, the next most likely progression would be to F major, followed in probability by A minor. In this sense, the similarity that any chord will follow another is not a rigorous probability in the mathematical sense but rather a general concept of the frequency of certain chord progressions in particular genres of music. As would be understood by one skilled in the art having the present description, drawings and claims before them, when the lead track results in the calculation of a different chord, then the chord flower (2750) will change. For example, let's say the next partition of the main music track is currently determined to correspond to B
<img file="MX345588B_D0081.tif" />
124
<img file="MX345588B_D0082.tif" />
ΙΝΤΙΤΤυΤΟ MEXICAN
DE Ι, Α PaOHEAD INDUSTRIAL flat major, then the center of the flower ItRSS'erary a capital ΕΓ with a flat symbol. In turn, the other chord found in the key of C major would rotate around the B flat in an arrangement that indicates the relative probability that any particular chord is next in the progression.
Track sharing module
Returning to the diagram of the system (100) in Figure 1A, the track-sharing module (148) can allow the transmission and reception of tracks or multi-track recordings for the system (100). In one embodiment, said tracks can be transferred to or received from a remote device or server. The track sharing module (148) may also perform administrative operations related to track sharing, such as allowing access to an account and exchanging payment and invoice information.
Sound finder module
The sound finder module 150, also shown in FIG. 1A, may implement operations related to finding a previously recorded track or a multi-track recording. For example, based on an audible input, the sound finder module 150 can search for tracks and / or similar multi-track recordings that were previously recorded. This search can be executed in a
125 A Aí P f ot LA P »OPIEO<sub>TO</sub>, A V'A E5 wwm<sub>w</sub>'particular device (50) or other devices or servers on the network. The results of this search can then be presented through the device and a track or multi-track recording and these can be subsequently had, purchased or otherwise acquired for use on the device (50) or otherwise. within the system (100).
Genre comparator module
Genre comparator module 152, also shown in FIG. 1A, is configured to identify chord sequences and beat profiles that are common to a musical genre. That is, a user can enter or select a particular genre of music or an example band that has a genre associated with the genre comparer module (152). Processing for each recorded track can be performed by applying one or more characteristics of the indicated genre to each generated audio track. For example, if a user indicates jazz as a desired genre, the quantization of a recorded audible input may be applied such that the timing of the beats may tend to be syncopated. Also, the resulting chords generated from the audible input can comprise one or more chords that are traditionally associated with jazz music. Also, the number of blue notes can be 25 more than what would be allowed in a classical piece.
126
<img file="MX345588B_D0083.tif" />
<img file="MX345588B_D0084.tif" />
Chord Comparator Module ————
The chord comparator module (154) provides pitch and chord related services. For example, the chord comparator module (154) can perform intelligent pitch correction on a monophonic track. Such a track may be derived from an audible input and pitch correction may include modifying a frequency of the input to align the pitch of the audible input with a particular predetermined frequency. The chord comparator module (154) can also build and refine an accompaniment for an existing melody included in a previously recorded multi-track recording.
In one embodiment, the chord comparer module (154) may also be configured to dynamically identify the probability of future chords appropriate for an audio track based on previously played chords. In particular, the chord comparer module 142) may, in one embodiment, include a music database. Using a hidden Markov model in conjunction with this database, the probabilities for a future chord progression can be determined based on the previous chords occurring on the audio track.
<img file="MX345588B_D0085.tif" />
127
<img file="MX345588B_D0086.tif" />
ΐΝ'τιη γ, τ, miicAt ·· :. <sup>;</sup>><sup>F</sup> .A ΜΟΛ'ϊι; αΪ
Network environment ______
As explained above, the device (50) can be. any device capable of executing the processes described above, and does not need to be in a network with any other device. However, Figure 28 shows the components of a potential embodiment of a network environment in which the invention can be practiced. Not all components may be required to perform the invention, and variations in the arrangement and type of components may be made without departing from the spirit and focus of the invention.
As shown, the system (2800) of Figure 28 includes local network areas (LANs) / wide area networks (WANs) - (network) (2806), wireless network (2810), client devices (2801-2805), Music Network Device (MND) (2808), and peripheral input / output devices (2811-2813). One or more of the client devices (2801-2805) may be comprised of a device (100) as previously described. Of course, while various examples of client devices are illustrated, it should be understood that, in the context of the network depicted in Figure 28, client devices 2801-2805 can include virtually any computing device capable of processing signals. audio and send related data
128
<img file="MX345588B_D0087.tif" />
with audio over the network, such as network ...... (2.8Q6), wireless network (2810), or the like. Client devices (2803-2805) can also include devices that are configured to be portable. Thus, client devices (2803-2805) can include virtually any portable computing device capable of connecting to another computing device and receiving information. Such devices include portable devices such as cell phones, smart phones, display pagers, radio frequency (RF) devices, infrared (IR) devices, personal digital assistants (PDAs), computers. laptops, laptop computers, notebook computers, tablet computers, embedded devices combining one or more of the above devices, and the like. And this is how client devices (2803-2805) differ widely in terms of capabilities and features. For example, a cell phone may have a numeric keypad and a low-line monochrome LCD screen on which only text can be displayed. In another example, a network-enabled mobile device may have a multi-touch screen, touch pen, multi-line color LCD screen on which both text and graphics can be displayed.
129
<img file="MX345588B_D0088.tif" />
Client devices (2801-2805) can also include virtually any computing device capable of communicating over a network to send and receive information, including track information and social media information, by executing audibly generated track search queries, or similar. The array of such devices may include devices that are typically connected using a wired or wireless communications medium such as personal computers, multiprocessor systems, consumer-programmable or microprocessor-based electronics, networked personal computers, or the like. In one embodiment, at least some of the client devices (2803-2805) can operate over a wired and / or wireless network.
A network-enabled client device can also include a browser application that is configured to receive and send web pages, network-based messages, and the like. Browser application 20 can be configured to receive and display graphics, text, multimedia, and the like, employing virtually any web-based language, including Wireless Application Protocol (WAP) messages and the like. In one embodiment, the browser application is enabled to use Markup Language
130 for portable devices (HDML, for its
<img file="MX345588B_D0089.tif" />
English), Wireless Markup Language (WML), WMLScript, JavaScript, Standard Generalized Markup Language 25 (SMGL), b
Hypertext Markup Language (HTML), Extensible Markup Language (XML) and the like, to display and deliver diverse content. In one embodiment, a user of the client device may employ the browser application to interact with a client that uses messages, such as a text message client, an email client, or the like, to send and / or receive messages. .
Client devices (2801-2805) may also include at least one other client application that is configured to receive content from another computing device. The client application may include an ability to provide and receive textual content, graphic content, audio content, and the like. The client application may also provide self-identifying information, including a type, capacity, name, and the like. In one embodiment, client devices (3001-3005) can uniquely identify themselves through a variety of mechanisms, including a phone number, a Mobile Identification Number (MIN), an electronic serial number
131
„.ΙΡΙ 0¾ ΙΝδΤΠΊΗΏ MKXICANC» ^^ '** £ 1 * 5 **
INDUSTRIAL PROPERTY (ESN) or other mobile device identifier. The information may also indicate the format of the content that the mobile device is enabled to use. Said information can be provided in a network packet, or the like, sent to the MND (108) or other computing devices.
The client devices (2801-2805) can also be configured to include a client application that allows the end user to log into a user account that can be managed by another computing device, such as the MND (2808), or similar. . Such a user account, for example, may be configured to allow the end user to participate in one or more social media activities, such as submitting a track or multi-track recording, searching for tracks or similar recordings to an audible input, downloading a track or recording and participate in an online music community, particularly one centered around sharing, reviewing and discussing the tracks and multi-track recordings produced. However, participation in various social media activities can also be done without logging into the user account.
In one embodiment, a musical input comprising the melody can be received by client devices (2801-2805) over the network (2806) or (2810) of MND (3008) from any other processor-based device with the
WICKED
MEXICAN INSTITUTE
OF THE fROFIECAP. ,,, INDUSTRIAL ------- ability to transmit said musical input. The musical input containing the melody can be pre-recorded or captured live by the MND (2808) or other processor-based device. Additionally or alternatively, the melody can be captured in real time by the client devices (2801-2805). For example, a melody generating device can generate a melody and a microphone in communication with one of the client devices (2801-2805) can capture the generated melody. If the musical input is captured live, the system typically searches at least one bar of music before the musical key and melody chords are calculated. This is analogous to musicians playing in a band, where an accompanying musician can typically listen to at least one measure of the melody to determine the musical key and chords being played before contributing any additional music.
In one embodiment, the musician can interact with client devices (2801-2805) in order to accompany a melody, treating a client device as a virtual instrument. Additionally or alternatively, the musician accompanying the melody may sing and / or play a musical instrument, such as an instrument played by the musician, to accompany a melody.
<img file="MX345588B_D0090.tif" />
133
<img file="MX345588B_D0091.tif" />
MEXICAN INSTITUTE • E LA INDUSTRIAL PROPERTY
The wireless network (2810) is configured to couple client devices (2803-2805) and their components with the network (2806). The wireless network (2810) may include a variety of wireless sub-networks that may later overlap independent ad hoc networks and the like, to provide an infrastructure-oriented connection for client devices (2803-2805). Such sub-networks can include mesh networks, wireless LANs (WLANs), cellular networks, and the like. The wireless network 2810 may further include an autonomous system of terminals, gateways, routers, and the like connected by wireless radio links and the like. These connectors can be configured to move freely and randomly and arrange themselves arbitrarily, so that the topology of the wireless network 2810 can change rapidly.
The wireless network (2810) may further employ a plurality of access technologies including 2 ^ (2G), 3 & (3G), 4 & (4G) generations of radio accesses for cellular, WLAN, mesh wireless router (WR) systems. , for its acronym in English) and the like. Future 2G, 3G, 4G access technologies and access networks may allow wide area coverage for mobile devices, such as client devices (2803-2805) with various 25 degrees of mobility. For example, the wireless network (2810)
<img file="MX345588B_D0092.tif" />
134
IMPI
MEXICAN INSTITUTE DE LA FRORIEDAO INDUSTRIAL can allow a radio connection through a radio network access such as Global System for Mobile Communication (GSM), General Radio Packet Services (GPRS, for its acronym in English) in English), GSM Extended Data Environment (EDGE),
Broadband Code Division Multiple Access (WCDMA) and the like. In essence, the wireless network (2810) can include virtually any communication mechanism by which information can travel between client devices (2803-2805) and another computing device, network, or the like.
The network (2806) is configured to couple network devices with other computing devices, including the MND (2808), client devices (2801-2802) and via the wireless network (2810) to client devices (2803 -2805). The network (2806) is enabled to employ any form of computer-readable media to communicate information from one electronic device to another. Also network 106 can include the Internet in addition to local area networks (LANs), wide area networks (WANs), direct connections, such as through a serial bus port (USB), other forms of computer-readable media, or any combination thereof. In an interconnected set of LANs, including those based on different architectures and protocols, a router acts
135
<img file="MX345588B_D0093.tif" />
as a link between LANs, allowing messages to be sent from one to another. In addition, communication links within LANs typically include twisted pair cable or coaxial cable, while inter-network communication links can use analog telephone lines, fully or fractionally dedicated digital lines including Ti, T2, T3 and T4 , Integrated Services Digital Networks (ISDNs), Digital Subscriber Lines (DSLs), wireless links including satellite links or other communication links known to those skilled in the art. On the other hand, remote computers and other related electronic devices could be remotely connected to either LANs or WANs via a modem and a temporary telephone link. Basically, network 2806 includes any communication method by which information can travel between computing devices.
In one embodiment, client devices (28012805) can communicate directly, for example, using a port-to-port configuration.
Additionally, the communication media typically includes computer-readable instructions, data structures, program modules, or other transport mechanisms, and includes any delivery media information. As an example, the media
MEXICAN INSTITUTE PE IA RROWEDAD INDUSTRIAL include wired media such as twisted pair, coaxial cable, fiber optics, waveguides, and other wired media, and wireless media such as acoustic, RF, infrared, and other wireless media.
Various peripherals including 1/0 devices (2811-2813) can be attached to client devices (2801-2805). The multi-touch pressure tablet (2813) can receive physical inputs from a user and distribute them as a USB peripheral, although not limited to USB and 10 other interface protocols can also be used, including but not limited to ZIGBEE, BLUETOOTH or the like. . Data transported via an external device and the pressure tablet 2813 interface protocol may include, for example, data in MIDI format, although data 15 or other formats may be transmitted over this connection as well. A similar pressure tablet (2809) may alternatively be physically integrated with a client device, such as a mobile device (2805). A headset (2812) can be connected to an audio port or other wired or wireless I / O interface of a client device, providing an example arrangement for a user to listen to a looped playback of a recorded track, along with other audible outputs of the system. The microphone (2811) can be connected to client devices (2801-2805) through the<sup>13</sup> IMPIAS
MEXICAN INSTITUTE η-rBCEiEDAL {V,
- ·> «»;.: «Ι ·. · ΊΐΑΐ. Re-entry of audio or another connection. Alternatively, or in addition to the headband (2812) and the microphone (2811), one or more other speakers and / or microphones can be integrated to one or more client devices (2801-2805) or other peripheral devices (2811-2813 ). Also an external device may be connected to pressure tablet (2813) and / or client devices (101-105) to provide an external source of sound samples, waveforms, signals, or other musical inputs that can be played by external control. Said external device may be a MIDI device to which a client device (2803) and / or pressure pad (2813) can route MIDI events or other data in order to trigger audio playback from an external device (2814). However, formats other than MIDI can be used by such an external device.
Figure 30 shows an embodiment of a network device (3000), according to one embodiment. The network device (3000) may include many more or fewer components than those shown. The components shown, however, are sufficient to describe an illustrative embodiment for practicing the invention. The network device 3000 may represent, for example, MND 2808 of FIG. 28. Briefly, the network device (3000) can include any computing device capable of
138
<img file="MX345588B_D0094.tif" />
connecting to the network (2806) to allow a user to send and receive tracks and track information between different accounts. In one embodiment, said track distribution, or sharing, is also executed between different client devices, which can be managed by different users, system administrators, business entities, or the like. Additionally or alternatively, the network device (3000) may allow sharing of a tune, including melody and harmony, produced with client devices (2801-2805). In one embodiment, said melody or tune-tone distribution, or sharing, is also executed between different client devices, which can be managed by different users, system administrators, business entities, or the like. In one embodiment, the network device (3000) also operates to provide a better musical key and / or similar chord for a melody from a collection of musical keys and / or chords.
Devices that can operate as network devices (3000) include different network devices, including but not limited to personal computers, desktop computers, multiprocessing systems, consumer-programmable or microprocessor-based electronics, personal computer networks, servers , network applications and the like. as the picture shows
<img file="MX345588B_D0095.tif" />
30, the network device (3000) includes the processing unit (3012), the video display adapter (3014), a mass memory, all in communication with each other via a bus (3022). Mass memory generally includes RAM (3016), ROM (3032), and one or more permanent mass storage devices, such as a hard drive (3028), tape drive, optical drive, and / or floppy drive. The mass memory stores the operating system (3020) to control the operation of the network device (3000).
Any general purpose operating system can be used. The basic input / output system (BIOS) (3018) is also provided to control the low level of operation of the network device (3000). As illustrated in Figure 30, the network device (3000) can also communicate with the Internet, or with some other communications network, via the network interface unit (3010), which is constructed for use with various communication protocols including the TCP / IP protocol. The network interface unit (3010) is known to 20 times as a transceiver, transceiver device, or a network interface card (NIC).
Mass memory is described above and illustrates another type of computer-readable media, specifically called computer-readable storage media.
Computer-readable storage media can
140
<img file="MX345588B_D0096.tif" />
include volatile, non-volatile, removable and non-volatile media
- lili · - 11 removable devices implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules or other data. Examples of computer-readable storage media include RAM, ROM, EEPROM, flash memory. Or other memory technology, CD-ROM, digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic storage disc, and other magnetic storage devices, or any other means that can be used to store the desired information and which can be accessed through a computing device.
As shown, the data repositories (3052) may include a database, text, spreadsheet, folder, file, or the like, which can be configured to maintain and store user account identifiers, email addresses, IM addresses, and / or other network addresses; group identifier information;
tracks or multi-track recordings associated with each user account; rules for sharing tracks and / or recordings, billing information; or similar. In one embodiment, at least some of the data repositories (3052) may also be stored in other components of the
<img file="MX345588B_D0097.tif" />
<img file="MX345588B_D0098.tif" />
network devices (3000), including but not limited to lim + dna. cd-rom / dvd-rom (3026), hard disk (3028), or similar.
Mass memory also stores program code and data. One or more applications (3050) are loaded into mass memory and run on top of the operating system (3020). Examples of application programs may include transcoders, scheduler, calendars, database programs, word processing programs, HTTP programs, customizable user interface programs, IPSec applications, encryption programs, security programs, message servers. SMS, IM message servers, email servers, account managers, and so on. The network server (3057) and Music Service (3056) can also be included as application programs within the applications (3050).
The network server (3057) represents a variety of services that are configured to provide content, including messages, over a network to another computing device. Thus, the network server 3057 includes, for example, a network server, a file transfer protocol (FTP) server, a database server, a content server, or the like. . The res server (3057) can provide the content including messages over the network using a
142
IMPIOUS
MEXICAN INSTITUTE LA MONEDAD variety of formats, including, but not limited to<sup>l</sup>or<sup>n</sup>''3'<sup>To the</sup>WAP,
HDML, WML, SMGL, HTML, XML, cHTML, xHTML, or simi-laj.eb. In one embodiment, the network server 3057 can be configured to allow a user to access and manage user accounts and shared tracks and multi-track recordings.
The music service (3056) may provide various functions related to enabling an online music community and may further include a music comparer (3054), a rights manager (3058) and tune data. The music comparator (3054) can match similar tracks and multi-track recordings including those stored in the data stores (3052). In one embodiment, such a comparison may be requested by the sound finder or the MTAC on a client device which may, for example, provide audible input, tracks, or multi-track recordings to be compared. The rights manager (3058) allows a user associated with an account to upload tracks and multi-track recordings. Such tracks and multi-track recordings can be stored in one or more data repositories (3052). The rights manager (3058) may further enable a user to provide controls for the distribution of the provided tracks and multi-track recordings, such as restrictions based on a
143
<img file="MX345588B_D0099.tif" />
'·<sup>Α</sup>,^<sup>θΜί</sup>^ Γ ÍNDUSTUML relationship or membership in an online music community, payment, or intended use of a track or multi-track recording. Using the rights manager (3058), a user can also restrict all access to the rights of a stored track or multi-track recording, thus allowing an unfinished recording and other work in progress to be stored without the community. can review before the user considers it ready.
The music service (3056) may also host or otherwise enable solo or multiplayer games to be played by and between various members of the online music community. For example, a multi-user role playing the game hosted by the Music Service (3056) can be configured in the music recording industry. Users can select a role for their character that is typical of the industry. The game user can then progress their character through music creation using their client device (50) and, for example, RSLL (142), and MTAC (144).
The message server (3056) can include virtually any computing component or components configured and arranged to forward messages from message user agents, and / or other message servers, or deliver messages. Thus, the message server (3056) can
144
<img file="MX345588B_D0100.tif" />
X ΑΨΑ AA í
MEXICAN INSTITUTE. . . J j. JJJ __OE LA MUFIEPAU include a menswje® transfer manager<sup>TO THE</sup> p communicate a message using a variety of messaging protocols, including, but not limited to, SMS, IM, MMS, IRC, RSS, mIRC, a variety of text message protocols, or a variety of other types of messages. In one embodiment, the messaging server 3056 may allow users to initiate and / or otherwise conduct chat sessions, VOIP sessions, text message sessions, or the like.
It is noted that while the network device 3000 is illustrated as a single network device, the invention is not that limited. For example, in another embodiment, a music service, or the like, the network device (3000) may reside on one network device, while an associated data repository may reside on another network device. In yet another embodiment, various music and / or message components for forwarding may reside on one or more client devices, operate in a port-to-port configuration, or the like.
I Gaming environment
To further facilitate music creation and composition, Figures 31-37 illustrate one embodiment in which a game interface is provided as the user interface for the previously described music compilation tools. In this way, it is believed that
<img file="MX345588B_D0101.tif" />
INSTITUTE NtaUCA · '*<sup>1</sup> FROM THE «INDUSTRIAL ONUJAI friendliest with interference with end. As will be
145 user interface will be less intimidating, the user so as to minimize any apparent creative musical process of a user from the following description, the interface "for games provides visual cues and cues associated with one or more functional aspects described above with the in order to simplify, optimize and incentivize the music compilation process. This allows end users (also referred to in this regard as players) to use professional quality tools to create professional quality music without requiring those users to have any experience in music theory or in the operation of music creation tools. .
Returning to FIG. 31, an exemplary embodiment of a first interface screen (3100) is provided. In this interface, the player can be provided with a studio view from the perspective of a music producer sitting behind a mixing console. In the embodiment of Figure 31, three different studio rooms are displayed in the background: a room for the lead instrument / voice (3102), a percussion room (3104), and a back room (3106). As would be understood by those skilled in the art having the present description, drawings and claims before them, the number of rooms could be
1 «WICKED
INSTITUTO-MEXICANO DE LA PROPERTY J industrial> ^^^ 251 larger or smaller, the functionality provided in each room can be differentially subdivided and / or additional options can be provided in the rooms. Each of the three rooms illustrated in Figure 31 may “include one or more musician avatars that provide visual cues that illustrate the nature and / or purpose of the room, as well as further provide cues for genre, style, and / or nuanced performance of the music performed by the avatars and the variety of instruments being used. For example, in the embodiment illustrated in Figure 31, the lead voice / instrument room (3102) includes a female pop singer, the back room (3104) includes a rock drummer, and the back room (3106) It includes a country violinist, a rock bassist, and a hip-hop keyboardist. As will be explained in greater detail later, the selection of musician avatars, in conjunction with other aspects of the gaming environment interface, provides an easy-to-understand visual interface, through which different tools described above can be easily implemented by the user. most novice of end users.
To start creating music, the player can select one of these rooms. In one embodiment, the user can simply select the room directly<sup>147</sup> IΜ ΡI
MEXICAN INSTITUTE <sup>V </sup>DE LA PROFIEDAÍ IN DUST11A L using a mouse and other input device.
Alternatively, one or more buttons can be provided to correspond to different study rooms. For example, in the mode illustrated in Figure 31, selection of a main room button (3110) will transfer the player to the main voice / instrument room (3102), selection of a percussion room button (3108) will transfer the player to the percussion room (3104); and selecting an accompaniment room button (3112) will transfer the player to the accompaniment room (3106).
Other selectable buttons may also be provided, as shown in Figure 31. For example, a record button (3116) and a stop button (3118) may be provided to start and stop the recording of any music made by the user. final in the studio room (3100) by the live loop replay recording session module (142) (see Figure 1A). A settings button (3120) may be provided to allow the player to alter various settings, such as the desired genre, tempo, rhythm and volume, and so on. A search button (3122) may be provided to allow a user to initiate the sound search module (150). Buttons for saving (3124) and deleting (3126) the player's music composition may also be provided.
148
<img file="MX345588B_D0102.tif" />
<img file="MX345588B_D0103.tif" />
Figure 32 presents an example embodiment of a. quarter instrument / lead voice (3102). In this embodiment, the interface for this studio room has been configured to allow an end user to create and record one or more main vocal and / or instrument tracks for a musical compilation. The main voice / instrument room 3102 may include a control space 3202 similar to that described above in conjunction with Figures 12-13. Thus, as described above, control space 3202 may include a plurality of partition indicators 3204 to identify each of the partitions (eg, musical measures) on the track; vertical lines (3206) illustrate the pulse within each measure, horizontal lines (3208) identify the various fundamental frequencies associated with a selected instrument (such as a guitar indicated by the instrument selector (3214) (shown in figure 32) , and a play bar identifies the specific part of a live loop that is currently playing.
In an example illustrated in Figure 32, the interface illustrates the audio waveform (3210) of a track that has already been recorded, presumably earlier in the player session, however the user can also connect enrichment assisted ( I polished up) audio tracks
149 pre-existing particularly together
<img file="MX345588B_D0104.tif" />
Mexican INSTITUTE OF INDUSTRIAL PROPERTY with the sound search module (150) (as called by the search button (3122) (see figure 31). In the example illustrated in figure 32, the audio form wave ( 3210) has also been converted into its note morphology (3212) in correspondence to the fundamental frequencies of a guitar, as indicated by the instrument selector (3214). As should be understood, by using various instrument selector icons that can be dragged over the control space (3202), the player may be able to select another of one or more instruments, which would cause the audio waveform to original is converted to a different morphology of notes corresponding to the fundamental frequencies of the newly or additionally selected instrument (s). The player can also alter the number of bars, or the number of beats per measure, which can also cause the audio waveform to be quantized (by quantizer (206) (see figure 2)) and aligned in the time with newly altered ^ 0 timing. It should also be understood that while the player may select to convert the audio waveform into a morphology of notes associated with an instrument, the player does not have to do so, thus allowing one or more original sounds from the audible input to be included from
150
IMPI
NSUTUTO (AEXiCANC OF THE INDUSTUAL raowitJAf
<img file="MX345588B_D0105.tif" />
substantial way on the audio track <sup>gi1</sup> original timbre.
As shown in Figure 32, an avatar of a singer (3220) can also be provided in the background. In one embodiment, this avatar can provide an easily understandable visual indication of a specific musical genre that has been previously defined in the genre comparator module (152). For example, in Figure 32, the singer is illustrated as a pop singer. In this case, the processing of the recorded track 3210 can be performed by applying one or more features associated with pop music. In other examples, the singer may be illustrated as an adult male, a young man or a girl, a barber-shop quartet, as an opera or opera diva.
Broadway, a country music star, a hiphop musician, a British Invasion rock singer, a folk music singer, among others with resulting tone, rhythms, modes, musical textures, timbres, expressive qualities, harmonies, etc. that people commonly understand that they are associated with each type of singer. In one embodiment, to provide additional entertainment value, the singer's avatar 3220 can be programmed to dance and any other way to act as the avatar that is involved in a recording session. <sup>151</sup> IMPIOS »NSTITUTO M EX! CA NO f) í LA PROriEHAO 'OtesíaJ INDUSTRIAL recording perhaps even in sync with the musical track.
The main voice / instrument room interface (3102) may further include a track selector (3216). The track selector (3216) allows a user to record or create multiple main takes and the selection of one or more of the takes to be included within the music compilation. For example, in Figure 32, three track windows, labeled 1, 2, and 3 are illustrated, each showing a miniature representation of an audio waveform from the corresponding track in order to provide a visual cue for the audio associated with each track. The track in each track window can represent a separately recorded audio take. However, it should be understood that copies of an audio track can be created, in which case each track window can represent different instances of a single audio waveform. For example, track window 1 can represent an unaltered vocal version of the audio waveform, track window 2 can represent the audio waveform as converted to a note morphology associated with a guitar, and window Track 3 could represent the same audio waveform as converted to a note morphology associated with a piano. As would be understood by those skilled in the art having the present description,<sup>152</sup> iMPI «. 'NITITUTOMIXICAN
DE LA PSOFIEDAU ΛΖ 'industrial S55! drawings and claims before them, there is no need for a particular limitation on the number of tracks that can be performed by the track selector (3216).
A track selection window (3218) is provided to allow the player to select one or more of the tracks for inclusion in the musical compilation by, for example, selecting and dragging one or more of the three track windows to the selection window. (3218). In one embodiment, the selection window (3218) can also be used to engage the MTAC module (144) in order to generate the best shot from the multiple shots 1! 2 and 3.
The main voice / instrument room (3102) may also include a plurality of buttons to allow one or more functions associated with the creation of a main vocal or instrumental track. For example, a minimize button (3222) may be provided to allow a user to minimize the grid (3202); the sound button (3224) may be provided to allow a user to mute or unmute the sound associated with one or more audio tracks, a solo button (3226) may be provided to mute any accompanying audio that has been generated through the system (100) based on the audio waveform (3210) or its morphology in order to allow the player to concentrate on aspects associated with the audio
Hee mexica institute:
FROM THE MAIN INDUSTRIAL PROPERTY, a new track button (3228) may be provided to allow the user to start recording a new main track; A morph button (3230) activates frequency detector and shifter (208) and (210) operations on the audio waveform in a control space (3202). A set of buttons can also be provided to set a reference pitch to help provide a vocal track. Thus, the pitch switch button (3232) can enable and disable a reference tone, the pitch up button (3234) can increase the frequency or reference pitch, the pitch down button ( 3236) can lower the pitch of the reference pitch.
Figure 33 illustrates an exemplary embodiment of a quarter percussion (3104). The interface for this room is configured to allow the player to create and record one or more percussion tracks for the musical compilation. The percussion room interface 3104 includes a control space similar to that described above in conjunction with FIG. 14. Accordingly, the control space may include a grid (3302) representing the playback and timing of separate sounds within one or more drum tracks, a playback bar (3304) to identify the specific part of the live loop that is currently playing. 25 being played at that time, and a plurality of partitions
154 IMPIOUS
INSTITUTO MlXlCA »» ./ of INDUSTaUL ^ 4-ϊχ (1-4) divided into multiple beats, with each box (3306) on the grid representing the time increments for the sounds associated with the related percussion instrument (where a unshaded box indicates that no h <sup>F</sup>5 sound associated with the related percussion instrument will be played at that time step, and a shaded box indicates that a sound associated with the timbre of the related percussion instrument will be played in that time step).
A drum segment selector (3308) may also be provided in order to allow a player to create and select multiple drum segments. In the example illustrated in Figure 33, only the partitions of a single drum segment A are shown. However, by selecting the drum segment selector 3308, additional segments can be created and identified as segments B, C, and so on. The player can then create different drum sequences within each partition of each different segment. The created segments can then be arranged in any order to create a more varied drum track for use in the music compilation. For example, a player may wish to create different drum tracks played repeatedly in the following order: A, A, B,
C, B, although any number of segments can be
<img file="MX345588B_D0106.tif" />
<img file="MX345588B_D0107.tif" />
<img file="MX345588B_D0108.tif" />
created and any order can be used. To facilitate multiple drum segment review and creation, a segment playback indicator (3310) may be provided to visually indicate the drum segment that is currently being played and / or edited, as well as the portion of the segment that is being played. and / or edited.
As can be seen in Figure 33, an avatar of a drummer 3320 can also be provided in the background. Similar to the performer avatar described in conjunction with the fourth voice / lead instrument (3102), the drummer avatar (3220) can provide an easily understandable visual indication of a musical genre and a specific playing style that corresponds to a genre that has been previously defined in a genre comparator module (152). For example, in figure 33, the drummer is illustrated as a rock drummer. In this case, the processing of the created percussion tracks can be performed for each percussion instrument by applying one or more of the previously defined percussion instrument features associated with rock music. In one embodiment, providing additional entertainment value the drummer's avatar (3320) can be programmed to dance and act in any way as the avatar is.
IMPI ^ iNs'TrnmtMtxicAK '.' ♦ 4 '| Λ Pt0lE0At involved in a recording session perhaps even in ~ sync with the music track. ........
The drum room interface 3104 may also include a plurality of buttons that allow one or more functions associated with the creation of one or more drum tracks. For example, the minimize button (3312) may be provided to allow a user to minimize the grid (3302), a sound button (3314) may be provided to allow a user to mute or unmute the sound associated with one or more audio tracks, A solo button (3316) may be provided to allow a user to toggle between mute and unmute the sound to stop playback of the other audio tracks so the player can focus on the drum track without distraction, the instrument button Percussion Additional (3318) adds an additional sub-track corresponding to a percussion instrument that can be selected by the player, and a swing button (3320) that allows the user to swing (ie syncopate) the notes.
Figures 34 AC present an exemplary embodiment of a back room interface (3106). The interface for this study room is configured to provide the player with a musical palette from which the user can select and create one or more backing tracks for a musical compilation. For example,
157 As shown in FIG. 34A, the player may be provided with an instrument class selector bar 3402 to allow the player to select an instrument class to accompany the main vocal and / or instrumental music tracks. In the illustrated embodiment, three classes are illustrated for selection - base (3404), keyboard (3406) and guitar (3408). As would be understood by one skilled in the art having the present description, drawings, and claims before them, any number of instrument classes can be provided including a variety of instruments, including brass, woods, and strings.
For illustrative purposes, let's say the player has selected the bass class 3404 in Figure 34A. In that case, the player is provided with an option to select from one or more avatars of musicians to play the accompanying instrument. For example, as shown in Figure 34B, the player may be provided with the option of selecting between a country musician (3410), a rock musician (3412), and a hip-hop musician (3414), which the Player I can select by clicking directly on the desired avatar. Of course, while three avatars are illustrated, the player can be allowed to select from one or more options. Arrows 3416 are also provided to allow the player to scroll through the avatar options, especially where more avatar options are provided.
After selecting a musician avatar in Figure 34B, the player can be provided with an option to select a specific instrument. For example, let's say the player has selected the country musician. As shown in Figure 34C, the player can be given the option of selecting between an electric bass (3418), a vertical bass (3420), or an acoustic bass (3422), which player 10 can select by clicking on the desired instrument.
The arrows (3424) are also provided to allow the player to scroll through the instrument options, which as would be understood by those skilled in the art having the present description, drawings and claims before them, may not be limited to just three types. low. Of course, while in the above sequence the instrument class is selected prior to selecting an avatar of a musician, it is contemplated that a player may be provided with the option of selecting an avatar of a musician prior to selecting an instrument class. .
Similarly, it is also contemplated that a player may be provided with the option of selecting a specific instrument prior to selecting an avatar of a musician.
<img file="MX345588B_D0109.tif" />
<img file="MX345588B_D0110.tif" />
<img file="MX345588B_D0111.tif" />
After the player has selected a musician, and instrument avatar, the system (100) creates an appropriate backing track by generating a set of backing notes based on one or more music tracks that are currently being played on the player. voice room / lead instrument (3102) (even if the other rooms have been muted), converting those notes to genre, timbre and musical style appropriate to the selected musician and instrument using the genre comparator module (152) and the harmonizer module (146) to harmonize one or more main tracks. Accordingly, a backing track for a specific instrument may have different sound, timing, harmony, blue note content and the like depending on the instrument and musician avatar selected by the player.
The backing room interface (3106) is also configured to allow the player to individually audition each of the multiple avatars of musicians and / or multiple instruments to aid in the selection of a preferred backing track. As such, once the musical instrument and avatar have been selected by the user and the corresponding backing track has been created as described above, the backing track is automatically
160
IMPI
NSTI TUTO MEXICANO DE LA PAOPIIOaí »tAJZSSjSSj INDUSTRIAL played in conjunction with other tracks created -nrey i amente .....
(lead, percussion, or backing) during live loop playback so that the player can, in virtually real time, evaluate whether the new track <sup>F</sup>5 accompaniment fits well. The player can then select to save the backing track, select a different musician avatar for the same instrument, select a different instrument for the same musician avatar, choose a completely new avatar and instrument, or delete the new backing track entirely. The player can also create multiple backing tracks by repeating the steps described above.
Figure 75 illustrates a potential embodiment of a graphical interface that represents the chord progression played in accompaniment to the main music. In one embodiment, this graphical user interface can be launched by pressing the flower button shown in Figures 34 A, 34 B and 34 C. In particular, this interface shows the chord progression that is generally forced on the multiple avatars accompanying in the accompaniment room (3106) subject to emitting any blue note (due to the gender and other aspects explained above together with figure 25) than the avatar can be built into its associated configuration file. Each avatar can have
<img file="MX345588B_D0112.tif" />
161
<img file="MX345588B_D0113.tif" />
NSTTTUTO MEXICANO OS LA FkCMXDAI?
INDUSTRIAL also certain arpeggio techniques (it's ripc-ir, amrrips ~ played note by note in a sequence) that are associated with the avatar due to the genre it belongs to or based on other attributes of the avatar. As illustrated in the faith example figure 35, the progression is G major, A minor, C major, A minor, with each chord played for the entirety of a partition according to the technique individually associated with each accompanying avatar in the accompanying room (3106). As would be understood by a person skilled in the art having the present description, drawings and claims before them, the chord progression may change multiple times within a single partition or the same chord may remain throughout a plurality of partitions.
Figure 36 illustrates an example interface by which a player can identify the portion of a musical composition that the player wishes to create or edit. For example, in the example interface shown in Figure 36, a tab structure (3600) is provided in which the player can select between an intro section, a stanza section, and a chorus section of a musical composition. . Of course, it should be understood that other portions of a musical composition are also available, such as a bridge, an ending, and the like. The portions that are available for editing in one
<img file="MX345588B_D0114.tif" />
Particular musical composition can be predetermined, manually selected by the player, or automatically set based on a selected musical genre. The order in which the different portions are ultimately arranged to form a musical composition can similarly be predetermined, manually selected by the player, or automatically set based on a selected musical genre. So, for example, if a novice user selects to create a pop song, the tab structure can be pre-populated with the expected elements of a pop composition, which generally includes an introduction, one or more verses, a chorus, a bridge and a conclusion. The end user can then be induced to create music associated with a first aspect of this general composition. After completing the first look of the overall composition, the end user can be directed to create another look. Each aspect can be scored individually and / or collectively to warn an end user if the tonality of the adjacent I elements is different. As would be understood by those skilled in the art having the present description, drawings and claims before them, using standard user interface manipulation techniques, portions of the composition can be removed, moved to other portions of the composition, the like.
163 copied and later
<img file="MX345588B_D0115.tif" />
modified
As shown in Figure 36, the tab for each portion of a music compilation may also include h "selectable icons to allow a player to identify and edit audio tracks associated with that portion, where a first row may illustrate the main track, the second row can illustrate the backing track, and the third row can illustrate the drum tracks. In the illustrated example, the intro section is shown including the main keyboard and guitar tracks (3602 and 3604, respectively); guitar, keyboard, and bass backing tracks (3606, 3608, and 3610, respectively); and a percussion track (3612). A chord selector icon (3614) may also be provided which, when selected, provides the player with an interface (such as in Figure 27 or Figure 35) allowing the player to alter the chords associated with the backing tracks. .
Figures 37A and 37B illustrate one embodiment of a file structure that can be provided for certain visual cues used in the graphical interface described above and stored in data repository (132). Returning to Figure 37A, a file (3700), also referred to herein as a musical resource, can be provided for each musician avatar that is
164
IMPI ^ j '"STITUTO MSXICANL" selectable by the player within the interface
For example, in Figure 37A, the top most .........._ top musical resource illustrated is for a hip-hop musician. In this embodiment, the musical resource can include visual attributes (3704) that identify the graphic image of the avatar that will be associated with the musical resource. The music resource may also include one or more functional attributes that are associated with the music resource and which, after selection of the music resource by the player, are applied to an audio track or compilation. The functional attributes can be stored within the music resource and / or provide a pointer or call to another file, object or process, such as the genre comparer (152). Functional attributes can be configured to affect a variety of settings or selection described above, including but not limited to the rhythm or tempo of a track, restrictions on the chords or keys to be used, restrictions on the instruments available, the nature of the transitions between notes, the i) structure or progression of a musical compilation, and so on.
In one embodiment, these functional resources may be based on the genre of music that could generally be associated with the visual representation of the musician. On occasions where visual attributes provide a representation of a specific musician, visual attributes
<img file="MX345588B_D0116.tif" />
<img file="MX345588B_D0117.tif" />
<img file="MX345588B_D0118.tif" />
Functional functions can also be based on the musical style of that particular musician.
Figure 37B illustrates another set of musical resources (3706) that can be associated with each selectable instrument, which can be a generic type of instrument (i.e. a guitar) or a specific brand and / or model of instrument (i.e. i.e. Fender Stratocaster, Rhodes Electric Piano, Wurlitzer Organ). Similar to the musical resources (3700) corresponding to the avatars of musicians, each musical resource (3706) for an instrument can include visual attributes (3708) that identify the graphic image of the instrument that will be associated with the musical resource, and one or more functional attributes (3710) of that instrument. As previously, the functional attributes 3710 can be configured to affect a variety of settings or selection described above. For an instrument, these can include the available fundamental frequencies, the nature of the transition between notes, and so on.
Using the graphical tools and game-based dynamics illustrated in Figures 31-37, a novice user will be more easily able to create more professional-sounding musical compositions that the user will be willing to share with other users for personal entertainment and even much entertainment. a very similar way in
166 IΜ ΡI
MEXICAN INSTITUTE; · de la eroriedai. industrial player can listen to commercially produced music. The graphic paradigm proposed in the context of a music authoring system in the present description would work equally well with respect to a variety of creative projects and endeavors that are generally executed by professionals due to the skill level that would otherwise be required. to produce even a poorly refined product it would be too high to be accessible to an ordinary person. However, by 10 simplifying routine tasks, even a novice user can be completing level projects with intuitive ease.
Processing Cache Memory
In one embodiment, the present invention can be implemented in the cloud where the systems and methods described above are used within a client-server paradigm. By downloading certain functions through a server, the processing power required for the client device is decreased. This increases both the number and type of devices in which the present invention can be implemented, allowing interaction with a mass audience. Of course, the extent for those functions that are executed by the server as opposed to the client can vary. For example, in one mode the server can be used to
167
<img file="MX345588B_D0119.tif" />
<img file="MX345588B_D0120.tif" />
<img file="MX345588B_D0121.tif" />
store and provide relevant audio samples<sub>r</sub> minT-ras p1 processing is executed on the client device. In an alternate mode, the server can store relevant audio samples and perform some processing before providing the audio to the client.
In one embodiment, client-side operations can also be executed through a single application that operates on the client device and is configured to communicate with the server. Alternatively, the user may be able to access the system and initiate communications with the server through an http browser (such as the Internet, Netscape, Chrome, Firefox, Safari, Opera, etc.). In some cases, this may require a browser plug-in to be installed.
In accordance with the present invention, certain aspects of the systems and methods can be implemented and / or improved through the use of an audio processing cache. More specifically, as will be described in more detail below, the processing cache allows for enhanced identification, processing, and retrieval of audio segments associated with identified or requested notes. As will be understood from the following description, the audio processing cache has a particular utility when systems and
<img file="MX345588B_D0122.tif" />
Methods described above are used with a client-server paradigm as previously described. In particular, in said paradigm, the audio processing memory would preferably be stored on the client side to improve latency and reduce server costs, although as explained later, the processing cache memory can also be stored remotely.
Preferably, the render cache is organized as an n-dimensional array, where n represents a number of attributes that are associated with, and are used to organize, the audio within the render cache. An exemplary embodiment of a processing cache 3800 in accordance with the present invention is illustrated in Figure 38. In this mode, the cache (3800) is organized as a four-dimensional array, where the four axes of the array represent (l) the type of instrument associated with a musical note, (2) the duration of the note, (3 ) the pitch and (4) the velocity of the note. Of course, other or additional attributes can also be used.
The instrument type can represent the corresponding MIDI channel, the pitch can represent an integer index of the respective semitone, the velocity can represent the force at which the note is played, and the duration can represent the length of the note in
169 ΐ ΪΜ ΡI *; <ΤίΤΌΤΟ WiiX'CAKC milliseconds. The entries (3802) in the cache<sup>-</sup>'-processing (3800) can be stored within ...... of — the —— array structure based on these four attributes, and each includes a pointer to allocated memory that contains the audio samples processed in the cache. Each cache entry can also include an indicator identifying a time associated with the entry, such as the time the entry was first written, the time it was last accessed, and / or the time that entry expires. This allows entries that were not accessed after a certain period of time has passed to be removed from the cache.
The render cache is also preferably kept up to a finite duration resolution, for example a 16th note, and is set in size in order to allow fast indexing.
Of course, other structures can also be used. For example, the render cache may be kept at a different finite resolution, or it may not be fixed in size if fast indexing is not necessary. Audio can also be identified using four more or less attributes, thus requiring an array that has more or fewer axes. For example, instead of a four-dimensional array, the inputs in Figure 38 may also be arranged as multiple arrays of three.
170
<img file="MX345588B_D0123.tif" />
dimensions, with a separate arrangement for each type of instrument.
It should also be understood that while an array is described as the preferred mode for caching, other memory conventions can be used as well. For example, in one embodiment, each audio input in the processing cache can be expressed as a hash value that is generated based on the values of the associated attributes. An example of a system that can be used to facilitate a cache system using this approach is Memcached. By expressing the audio in this way, the number of associated attributes can be increased or decreased without requiring significant changes to the associated code for searching and identifying cache entries.
Figure 39 illustrates an example of data flow using such a cache. As shown in Figure 39, process 3904 executes cache control. The process (3904) receives requests for a note from a customer (3902), and in response obtains a cached audio segment corresponding to that note. The note request can be any request for a specific note. For example, the requested note can be a note that has been identified by a user through any of the interfaces described above, a note identified by the module
INSTITUTE Mtxx »I> £ LA PROPISD. ' harmonizer, or any other source. Instead of identifying a specific grade, the grade request also identifies a plurality of attributes associated with a desired grade. Although this is generally referred to in the singular, it should be understood that a note request may involve a series or group of notes, which may be stored in a single cache entry.
In an example mode, the notes can be specified as MIDI 'note-on' data with a given duration of 10, while the audio is returned as a pulse code modulated (PCM) encoded audio sample. English). However, it should be understood that notes can be expressed using any attribute or attributes, and in any notation, including MIDI, XML, or the like. The retrieved audio sample can also be compressed or uncompressed.
As shown in FIG. 39, process 3904 communicates with process 3906, process 3908, and processing cache 3800. Process (3906) is / 0 configured to identify the required note attributes (such as instrument, note-on, duration, pitch, velocity, etc.) and process the corresponding audio using a library of available audio samples (3910 ). Audio processed by process (3906) 25 in response to a requested note is returned to process
172
IM FI MEXICAN INSTITUTE DE LA FROPiíOAD (3904), which provides the audio to the client ¿^^<sup>TO</sup>), and you can also write the processed audio to the render cache (3800). If a similar note is requested later, and the audio corresponding to that requested note is already available in the render cache, the process (3904) can retrieve the audio from the render cache (3800) without requesting that a new audio segment to be processed. In accordance with the present invention, and as will be described in more detail below, an audio sample can also be retrieved from the processing cache although it is not an exact match with the requested note. This recovered audio sample may be provided to process 3908, which reconstructs the note into one that is substantially similar to the audio sample that substantially corresponds to the requested note. As the process of recovering and rebuilding audio from the cache is generally faster than the process (3906) for processing new audio, this process significantly improves performance h
of the system. It should also be understood that each of the items illustrated in Figure 39, including processes (3904, 3906, and 3908), processing cache (3800), and sample library (3910) can be operated on the same device as the client, on a remote client's server, or on any other device; and that several of the
173 element can be distributed among 'ÍH'Wfso's' devices in a single mode. ------- Figure 40 describes an example method that can be used for processing notes requested by the faith.<sup>r</sup>5 cache control (3904). This example method is described assuming the use of a four-dimensional cache as illustrated in Figure 38. However, one skilled in the art having the present disclosure before them would be able to easily adapt the method for use with different cache structures.
At step 4002, a requested note is received from the customer 3902. In step 4004, it is determined whether the processing cache 3800 contains an entry that corresponds to the specific note requested. This can be accomplished by identifying the instrument with which the requested note is associated (i.e. a guitar, piano, saxophone, violin, etc.), as well as the duration, pitch, and velocity of the note, and subsequently determining if there is an input in cache that precisely matches each of these parameters. If there is, the audio is retrieved from the cache in step 4006 and provided to the client. If there is no exact match, the process proceeds to step (4008).
In step (4008), you determine if there is enough time to process a new audio sample for the note.
<img file="MX345588B_D0124.tif" />
INSTITUTO MEXICANA df la fwwím INDUSTRIA ·
<img file="MX345588B_D0125.tif" />
requested. For example, in one embodiment, the client may be configured to identify a specific time when the audio for the note should be provided. The time in which the audio must be provided can be a preset amount of time after the request was made. In an embodiment where live loop repetition is employed, as previously described, the time at which audio is provided may also be based on the time (or number of measures) until the end of loop 10 and / or or until the note will be played during the next loop.
In order to assess whether the audio can be provided within the time limit, an estimate of the amount of time to process and send the note is identified and compared to the specific time limit. This estimate can be based on numerous factors, including a predetermined estimate of the processing time required to generate the audio, the length of any backlogs or processing queues present at the time of the request, and / or the speed in the bandwidth of the connection between the client device and the device that provides the audio. To perform this step, it may also be preferable that the client's system clocks and the device on which the cache control 3904 is operating are synchronized. Whether
175
<img file="MX345588B_D0126.tif" />
determine that there is enough
MEXICAN INSTmrro. .____ _ „__ _“ «vtonsDA», r ~ .λ> · / the time to process then, in step (4016) the note is sent to the proucbu du, note processing (3906), where the audio stops the requested note is processed. Once processed, the audio can be cached 3800 in step 4018.
However, if it is determined that there is not enough time to process the note, then the process proceeds to step (4010). In step 4010, it is determined whether a near hit input is available. For the purposes of this description, a close hit is any note that is sufficiently similar to the requested note that it can be reconstructed, using one or more processing techniques, within an audio sample that is substantially similar to the audio sample that would be processed for the requested note. A close hit can be determined by comparing the instrument type, pitch, velocity, and / or length of the requested note with those of the notes already in the cache. Because different instruments are
If they behave differently, it should be understood that the range of inputs that can be considered a close hit will differ for each instrument.
In a preferred embodiment, a first search for a near hit entry may be made on a near cache entry along with the duration axis of the
<img file="MX345588B_D0127.tif" />
<img file="MX345588B_D0128.tif" />
<img file="MX345588B_D0129.tif" />
the same type even more with a greater being acceptable
176 cache memory, (that is, an entry with instrument, tuning and velocity). It is preferable that the search is for a duration input (within a certain range for the given instrument) than the requested note, since shortening the note often produces a better result than lengthening it. Otherwise, if there is no acceptable input along the duration axis, a second search may be to find a nearby cached entry along the tuning axis, that is, an input within a certain range of semitones.
In still another alternative, or if there are no acceptable entries on any of the duration or pitch axes, a third search for a nearby cached entry within a range along the velocity axis. The acceptable range at different speeds may, in some cases, depend on the specific software and algorithms used to perform the audio reconstruction. Many audio samples use various samples mapped to different velocity ranges for a note, as many real instruments have significant timbral differences in the sound produced depending on how loud the note is placed. Thus, it is preferable that a close hit along the velocity axis would be an audio sample that differs from the requested note only in amplitude.
177
<img file="MX345588B_D0130.tif" />
Still in another alternative,
ΙΜΡΪ
INSTITUTE M'XtCANC
Γ, Γ LA? ΜΜ> 1ΕΡΑΠ INDUSTRIAL or if there are no acceptable entries in the duration, tuning, or velocity axes a fourth search will be to find a nearby cached entry within a range along the instrument axis. Of course, it is understood that this strategy may be limited to only certain types of instruments that produce sounds similar to other instruments.
It should also be understood that while it is preferable to identify a close hit input that differs only by a single attribute (in order to limit the amount of processing required to reconstruct the audio sample), a close hit input can also be an input that differs. on two or more attributes of duration, tuning, speed, and / or instrument. Additionally, if multiple close hit inputs are available the audio sample to be used can be selected based on any one or more of a number of factors including, for example, the distance of the desired note in the
To array (for example, determining the shortest Euclidean distance in 'n' dimension space), the closest hash value based on the attribute, a weighting of the priority of each axis in the array (for example, audio that differs in audio is preferred over audio that differs in
<img file="MX345588B_D0131.tif" />
178
IMPI
MEXICAN INSTITUTE
INDUSTRIAL PROPERTY instrument) and / or the speed in the processing of the audio sample.
In another embodiment, close hits can be identified using a composite index approach. In this ^ 5 mode, each dimension in the cache is folded. To an approximation, this can be achieved by folding a certain number of bits from each dimension. For example, if the lower two bits of the tuning dimension are doubled, all tuning values can be mapped 10 to one of 32 values. Similarly the three bits below the duration dimension can be folded. As a result of this, all durations can be mapped onto one of 16 values. Other dimensions can be processed in a similar way. In another approach, a non-linear folding method can be used where the instrument dimension is assigned a similar sounding instrument with the same folded dimension value. The folded dimension values can be concatenated into a composite index, and the cached entries can be stored in a table that is ordered by the composite index. When a note is requested, relevant cached entries can be identified through a look-up based on the composite index. In this case, all results that match the composite index can be identified as 'close hit' inputs.
<sup>179</sup> IΜ PI
INSTITUTO MUlCANO Dt l> ^ kOFIF.DA »industrial 2
If, in step 4010, it is determined that a close hit input is available, the process proceeds to step 4012 where the close hit input is reconstructed (via note reconstruction process 3908) for. Fe generate an audio sample that corresponds substantially to the requested note. As shown in figure 40, the reconstruction can be performed in different ways. The techniques described below are provided as examples, and it should be understood that other reconstruction techniques may also be used. Furthermore, the techniques described below are generally known in the art for audio sampling and manipulation. Consequently, while the use of the techniques in conjunction with the present invention is described, the specific algorithms and functions to implement the techniques are not described in detail.
The rebuilding techniques described below can also be run on any device in the system. For example, in one embodiment, the ^ 0 rebuild techniques can be applied on a cache server or through a remote device attached to the cache server, where the rebuilt note is then provided to the client device. However, in another mode, the cached note itself can be transmitted to the client device, and the rebuild can
180
<img file="MX345588B_D0132.tif" />
IMPí
MEXICAN INSTITUTE OS LA RíCEISDAL INDUSTRIAL then be executed on the client. In this — the information identifying the note and / or instructions for executing the rebuild can also be transmitted to the client along with the cached note.
A Going back to the first technique, suppose that for example, the input close hit differentiated only in duration with the requested note. If the audio sample for the near hit is larger than what was requested, the audio sample can be reconstructed using a re-filtering technique where a new, shorter, filter is applied to the audio sample.
If the requested note is longer than the near hit input, the sustained note portion of the filter can be reduced to the desired length. Because the attack and decay are generally considered to be what gives an instrument its sonic character, manipulations on the sustained note can reduce the duration without significant impact on the color of the note. This is termed as filtering-reduction, h
Alternatively, a loop repetition technique can be applied. In this technique, instead of reducing the sustained note portion of an audio sample, a section of the sustained note section can be looped in order to lengthen the length of the note. However, it should be considered that randomly selecting a portion of the
181
IMPI
MSXICAN INSTITUTE
ΠΕ LA ηΟΗίΠΛΟ VSireF industrial ----— sustained note section for looping can result in docks and pops in the audio. In one embodiment, this can be remedied by crossfade from the end of one loop to the start of the next loop. In order to lessen any effects that may result from processing and adding various effects, it is also preferable that the cache input can be a raw sample, and that any additional digital signal processing is performed after the rendering has been completed. rebuild, for example, on the client device.
If the requested note is of a different pitch than the near hit input, the cached audio sample can be shifted in pitch to acquire the appropriate pitch. In one embodiment, this can be done in the frequency domain using FFT. In another embodiment, this can be done in the time domain using autocorrelation. In a scenario where the requested note is one octave higher or lower, the cached note can also be simply lowered or simply shortened to Ό acquire the proper pitch. This concept is similar to using a faster or slower tape player. That is, if the cached entry is shortened to play twice as fast, the pitch of the recorded material becomes twice as high, or an octave higher. If the cache entry is narrowed to play twice as slow, the pitch of the material
182 T lUf ΡI
Λ. .11 .Ιν,> * 3
MEXICAN INSTITUTE FC - * Τ ^ fX LA fHOMtOAL · V '-.!> -'. LJ®r IHCTJSTWAI.
engraving is divided in half, or one octave down.
Preferably, this technique is applied for cached entries that are approximately two semitones from the requested note, since narrowing or shortening an audio sample by more than that amount can cause an audio sample to lose its sonic character.
If the requested note is of a different velocity than the close hit input, the cached input may be offset in amplitude to match the new velocity. For example, if the requested note has a higher velocity, the width of the cache entry can be increased by the corresponding difference in velocity.
If the requested note is of a lower velocity, the amplitude of the cache entry can be decreased by the corresponding difference in velocity.
The requested note can also be of a different but similar instrument. For example, the requested note can be a specific note played on a heavy metal guitar, while the cache can only contain one note for a raw metal guitar. In this case, one or more DSP effects could be applied to the cached note in order to approximate a note on a heavy metal guitar.
After a close hit entry has been rebuilt using one or more of the techniques described
183
IMFI ^
INSTITUTO HUKM * · previously, it can be sent back to the citen ^ é ^ indication can also be provided to L- ^ usuer ^ -or ^ -bitch ”” informing the user that a reconstructed note has been provided. For example, in an interface such as the one shown in Figure 12A, suppose that note (1214) has been rebuilt. In order to inform the user that this note has been reconstructed from other audio, the note can be illustrated in a different way from the processed notes. For example, the reconstructed note can be illustrated in a different color from the other notes, such as a hollow note (as opposed to solid color), or any other type of indication. If the audio for the note is then processed later (as will be explained later), the visual representation of the note can be changed to indicate that a processed version of the audio has been received.
If, in step 4010, a cached near hit entry was not present, the closest available audio sample (as determined based on instrument, tuning, duration, and velocity attributes) can be retrieved. In one embodiment, this audio sample can be retrieved from cache (3800). Alternatively, the client device can also be configured to store in local memory a series of general notes to be used in circumstances when neither a processed note nor a close hit note is available.
184 IΜ ΡI
INSTITUTO MBXICANC ιτι IA HROI-IEDai '^ Tas ^ i reconstructed. Additional processing, such as that described above, can also be performed on this audio sample. A user interface on the client can also be configured to provide a visual indication to the user that an audio sample has been provided that is neither processed audio nor a reconstructed near beat.
At step 4016, a request is made to render note processing 3906 to process audio 10 for a requested note using sample library 3910. Once the note is processed, the audio is returned to the cache control (3904), which provides the processed audio to the client (3902), and writes the processed audio to the processing cache (3800) in step 15. (4018).
Figure 41 shows one embodiment of an architecture for implementing a processing cache with the present invention. As shown, a server (4012) is provided that includes an audio processing engine ^ 0 (4104) for processing audio as described above, and a cache-server (4106). The server (4102) may be configured to communicate with a plurality of different client devices (4108, 4110, and 4112) through a communication network (4118). The communication network (4118)
185
<img file="MX345588B_D0133.tif" />
It can be any network including the Internet, a cell phone network, Wi-Fi, and so on.
In an example embodiment shown in FIG. 41, device 4108 is a thick client, R device 4110 is a thin client, and device 4112 is a mobile client. A thick client, such as a full-featured desktop or laptop, typically has a large amount of memory available. As such, in one embodiment, the processing cache can be kept entirely on an internal hard drive of the thick client (illustrated as client cache 4114). A client is generally a device with less storage space than a client. Consequently, the processing cache for a thin client can be divided between the local hard disk (illustrated as the client cache (4116) and the server cache (4106)). In one embodiment, the most frequently used notes can be cached locally on the hard drive, while frequently used notes can be cached on the server. A mobile client (such as a cell phone or smartphone) generally has less memory than either a thick client or a thin client. Thus, the processing cache for a mobile client can be fully kept in the server cache (4106). Of course these are provided
<img file="MX345588B_D0134.tif" />
as examples and it should be understood that any of the above configurations may be used to
<img file="MX345588B_D0135.tif" />
<img file="MX345588B_D0136.tif" />
any type of client device.
Figure 42 shows another embodiment of an architecture for implementing a processing cache in accordance with the present invention. In this example, multiple Edge Cache Servers (4102-4106) can be provisioned and located to serve multiple geographic locations. Each client device (4108, 4110, and 4112) can then communicate with the edge cache server (4102, 4104, and 4106) closest to its geographic location in order to reduce the transmission time required to obtain an audio sample. cached. In this mode, if a client device requests audio for a note that was not previously cached on the client device, it determines whether the edge cache server includes either the audio for the requested note, or a close hit. for that note. If that happens, then the audio sample is obtained and / or reconstructed, respectively, and provided to the client. If such a cached entry is not available, the audio sample can be requested from the server (4102) which - according to the process described in association with figure 40 - can either provide a cached entry
187
<img file="MX345588B_D0137.tif" />
(either an exact match or a close hit) —Play the note.
Figure 43 illustrates one modality of the signal sequence between the client, the server, and an edge cache of Figure 42. Although Figure 43 refers to the client 4108 (for example, the complex client) and memory Edge Cache (4202), it should be understood that this signal sequence can be applied similarly to a thin client (4110) and (4112) and edge caches (4204) and (4206) in Figure 42. In Figure 43, the signal (4302) represents a communication between the server (4102) and the edge cache (4202). In particular, the server (4102) transmits audio data to the edge cache (4202) in order to send and preload the edge cache with audio content. This can happen either autonomously or in response to a request from a customer. Signal 4304 represents a request for audio content that is sent from client 4108 to server 4102. In one embodiment, this request can be formatted using the hypertext transfer protocol (http), although other languages or formats can also be used. In response to this request, the server (4102) sends a response back to the client, illustrated as a signal (4306). The answer signal (4306) provides the client (4108) with a
188
<img file="MX345588B_D0138.tif" />
redirect to a location in the location Tl ^ Ta ”cache (in the edge cache (4202), for example). The server (4102) may also provide a manifest that includes a reference to a cached content list. This list can identify all cached content, although preferably the list would only identify cached content that is relevant to the requested audio. For example, if the client (4108) requested violin audio for a central C, the server can identify all cached content for the violin notes.
The manifest can also include any encryption keys required to access the relevant cached content, as well as a time to live (TTL) that can be associated with each cached entry.
After receiving the response from the server (4102), the client (4108) sends a request (illustrated as signal (4310)) to the edge cache (4202) to identify the appropriate cache entry (for either the ^ 0 specified associated audio, or a close hit, etc.) based on the information in the manifest. Again, this request can be formatted using http, although other languages or formats can also be used. In one embodiment, the client (4108) executing the determination can also be executed remotely in the edge cache (4202). The
<img file="MX345588B_D0139.tif" />
signal (4310) represents the response from the server
<img file="MX345588B_D0140.tif" />
<img file="MX345588B_D0141.tif" />
Edge Cache Server for Client (4108) that includes the identified cache entry. However, if the request identified a cached entry that is beyond its TTL, or is otherwise unavailable, the response will include an indication that the request has failed. This may cause the client (4108) to retry this request to the server (4102). If the response (4310) contains the requested audio input, it can then be decrypted and / or decompressed, as needed, by the client (4108). If the cache entry was very close, it can also be rebuilt using the processes described above or their equivalents.
Figure 44 illustrates an alternate mode of signal sequence between client, server, and edge cache of the mode described in association with Figure 42. In this mode, the communication between the client (4108) and (4202) are similar to those described in figure 43 with the exception that, instead of the client (4108) contacting the server (4102) to obtain the memory location cache and a manifest of the cached content, the client (4108) directly sends the request for the audio content (4308) to the edge cache (4202).
190
<img file="MX345588B_D0142.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
Figures 45-47 illustrate three techniques that can be used to optimize the processes used to request and retrieve audio in response to a request from a customer. These techniques can be employed either on a server, R an edge cache, or any other device that stores and delivers audio content to the client in response to a requested note. These techniques can also be applied individually, or in conjunction with one another.
Turning first to Figure 45, an example method is described to enable a client to quickly and efficiently identify when there is insufficient time for audio to be provided from a remote server or cache. In block 4502, an audio request is generated on the client. The audio request can be a request for cached audio or a request for audio to be processed. A fault identification request, as well as a time when audio is required by the client (called as a time limit), can also * 0 be included with the audio request in block (4504).
The request failure may include an argument identifying whether to abort or proceed with the audio request if the audio cannot be provided to the client within the time limit. The timeout provided in the audio request is preferably a real time value. In this
<img file="MX345588B_D0143.tif" />
<img file="MX345588B_D0144.tif" />
<img file="MX345588B_D0145.tif" />
In this case, it is necessary for the client and the request to be synchronized in time. As would be understood by one skilled in the art having the present description, drawings and claims before them, other methods of identifying a time limit may also be employed. Preferably, the failure identification request and timeout are included in the header of the audio request, although they can be transmitted in any other portion of the request or as separate signals.
In block 4506, the audio request is transmitted from the client to the relevant cache or server. The server or cache receives the audio request at block 4508 and determines that the received audio request includes a block failure request 4510 received by the server or cache. In block 4512, the server or the receiving cache determines if the requested audio can be provided to the client in the time limit. This is preferably determined based on previously projected or determined times to identify and fetch the cached audio, process the note, and / or transmit the note back to the client. The time required to transmit the note back to the client may also be based on an identified latency time between the
<img file="MX345588B_D0146.tif" />
<img file="MX345588B_D0147.tif" />
<img file="MX345588B_D0148.tif" />
MEXICAN INSTITUTE DE LA ΡΕΟΜΪΟΛ »industmal
<img file="MX345588B_D0149.tif" />
transmission of the audio request and the received time.
If it is determined that the audio can be provided before the time limit, the audio is queued at block 4514, and the method for identifying, locating and / or processing the audio proceeds as described above. If it is determined that the audio cannot be provided before the time limit, a message is sent back to the client in block (4516) notifying the client that the audio will not be available at the time limit. In one embodiment, the notification can be transmitted as an http error message (412), although any other format can also be used. The client can take whatever actions are necessary at block 4518 to obtain and provide substitute audio. This can be accomplished by the client by identifying audio that is similar to that required by the requested note from a local cache, and / or by applying processing to previously stored or cached audio to approximate the requested note.
In block 4520, the server / cache checks if the request failure has identified whether to abort or continue if the event audio could not be provided within the time limit. If the request failure is set to abort, the audio request is discarded at block 4522 and no further action is taken. If the request failure fits
MEXICAN INSTITUTE DE LA PROPIEDAD INDUSTRIAL to continue, the audio request is placed Hpnhrn Hp ia rnia for processing in block (4514). In this case, the audio may be provided to the customer once completed and used to replace the substitute audio that has been faithfully obtained by the customer.
Figure 46 illustrates example processes for prioritizing audio requests in a queue. This process is particularly useful in conjunction with the implementation of the loop repeat live recording session 10 described above, as it is beneficial for any changes made by a user to a note in a live loop session which is desirable. that is implemented before the note is played during the next playback lap of the live loop. In block 4602, an audio request is generated by a client for a note that is used within a current live loop. The timing information related to the live loop is included in the audio request at block 4604. In one embodiment, the timing information can identify the duration of the h loop (referred to as the length of the loop). In another embodiment, the beat information may also include information identifying the position of the note within the loop (referred to as the note at the start time) as well as the current portion of the loop that is playing, 25 as can be identified by the position of a bar
194 τ Μ Ρ1 tΝ ^ ΤΓί'υΤΟ> * ί 1'KICA-HO '/ la ^' tOVUDAD V.
..HDICTRIAL ----- playback or playback head in the interface described „above (referred to as the playback time from the start). (An example mode of a live loop, and the relative timing information described in this paragraph is illustrated in Figure 48.)
Returning to FIG. 46, the audio request, along with the timing information, is sent to the server or cache at block 4606. In one embodiment, a timestamp indicating when the message has been sent can also be included with the message.
The audio request is received in block 4608 and a time to serve is determined in block 4610. For example, in one embodiment, if the audio request only includes information regarding the length of the loop, the service time can be calculated simply by dividing the length of the loop in half. This provides a statistical approximation of the length of time that is likely to be required before the live loop playback on the client reaches the location of the note for which the audio was requested.
In another embodiment, if the note start time and the playback time information from the start is included in the audio request, the service time can be more accurately calculated. For example, in this case, it can first be determined whether the start time of the
<img file="MX345588B_D0150.tif" />
195
ΙΜΡΪ
MEXICAN INSTITUTE ü! THE INDUSTRIAL PROHtTY note is greater than the playback time from the start (play head) (for example, the note was at a later position in the loop than the playback bar at the time the audio request was made). If the note start time R is longer, the service time can be calculated as follows:
service_time = note_start_time— production_time_from_start. If the play time from the start is greater than the note start time 10 (for example, the note was earlier in the loop than the play bar at the time the audio request was made) , the service time can be calculated as follows: service_time = (loop_length15production_time_from_start) + note_start_time. In another embodiment, the service time calculation may also include the addition of the projected latency time required for transmission of the audio data back to the client. The latency time can be determined by identifying the timestamp of when the audio request was sent, and calculating an identified elapsed time between the timestamp and the time the audio request was received by the server or memory. cache.
After the service time value is determined, the audio request is placed in a queue based on its service time. As a result, the audio requests with a shorter service time are R processed before those with a longer service time, thus increasing the probability that the audio requests will be processed before the next playback of the associated note in the live loop.
Figure 47 illustrates an example process for adding 10 repeating audio requests related to the same note.
In block 4702, an audio request is generated by a client. In block 4704, a track ID, a note ID, a start time, and an end time are included with the audio request. The track ID identifies the music track for which the audio request has been made, and the note ID identifies the note. Preferably, the track ID is a unique global ID, while the note ID is unique for each note within a track. The start time and end time identify the start and end positions of the note relative to the start of the track, respectively. In block 4706, the audio request and associated track ID, note ID, start time, and end time are transmitted to a server and / or cache.
197
<img file="MX345588B_D0151.tif" />
As shown in figure 47, in this mode<sub>F</sub> ai. server and / or cache has a queue (4720), which includes a plurality of track queues (4722). Each track queue (4722) includes a separate queue for processing the audio requests for an individual track. In block (4708), the server or cache receives the audio request and, in block (4710) identifies a queue of tracks (4722) in queue (4720) based on the track ID associated with the request for audio. Audio. At block 4712, the track queue is searched to identify any previously queued audio requests with the same note ID. If an audio request with the same ID is located, that request is removed from the track queue (4722) at block (4714).
The new audio request is then positioned within the respective one of the plurality of track queues (4722). This can be accomplished in one of several ways. Preferably, if a previous audio request with the same note ID has been located and discarded, the new audio request may replace the discarded request in the track ^ 0 queue (4720). Alternatively, in another embodiment, the new audio request can be positioned in the track queue based on the start time of the audio request. More specifically, notes with an earlier start time are placed earlier in the queue than notes with a later start time.
<img file="MX345588B_D0152.tif" />
<img file="MX345588B_D0153.tif" />
ΙΜΡΙ (^
INSTITUTE MSIC.CAHl
DH LA HIOPHL AD
As a result of the method described in ¿Scissors 4 ^ 7 ^, audio requests that are obsolete or Substituted are removed from the queue, thus conserving processing power. This is particularly useful when one or more users are making numerous successive changes to individual notes during a live looping session, as the ability of the system to quickly and efficiently process and deliver the most recently requested notes and avoid noise increases. processing notes that are no longer wanted or needed.
Effects chain processing
Figures 49-52 illustrate processes that can be used to apply a series of multiple effects to one or more music tracks based on the virtual musicians, instruments, and producer selected by a user to be associated with those music tracks, in particularly for the gaming environment described above. As should be understood from the following descriptions, by virtue of these processes, user-created tracks can be processed to better represent or mimic the styles, nuances, and trends of the available musicians, instruments, and producers that are represented in the environments. of games. As a result, a single track can sound significantly different based on the
199 selected musicians, instruments and producers », '' afer<sup>2 </sup>associated with the runway. _______
Referring first to FIG. 49, an example effect chain for applying R effects to one or more music tracks for a music compilation is illustrated. As shown, for each instrument track, a first set of effects (4902, 4904, and 4906) can be applied based on the avatar of the selected musician associated with that track. These effects are referred to in this document as musician role effects. A second series of effects (4904) can be applied to each of the instrument tracks based on the selected producer avatar. These are referred to in this document as effects corresponding to the role of producer. Although specific examples of applied effects will be described below, it should be understood that various effects can be used, and the number and order of effects that can be applied for each of the roles of musician and producer can be altered.
h
10 Figure 50 shows an example mode of the musician role effects that can be applied to a track. In this embodiment, a track (5002) is input to a distortion gear selection module (5004) that applies the relevant digital signal processing to the music track in order to substantially recreate the type of music.
200
<img file="MX345588B_D0154.tif" />
sound that can be associated with the real-life instrument represented by the selected virtual instrument through the game interface. For example, if track (5002) is a guitar track, one or more effects can be applied to the basic electric or acoustic guitar track (5002) in order to mimic and recreate the sound style of a particular guitar. including, for example, bridging, chorus, distortion, echo, envelopes, reverb, wah and even complex combinations of effects resulting in retro, metal, blues or grunge feel. In another example, effects can be automatically applied to basic electric keyboard tracks 5002 to mimic keyboard types such as a Rhodes Piano or a Wurlitzer Electric Organ. If track 5002 is a basic drum track, a pre-configured drum sound kit can be applied through the effects chain based on a selected set of drums. Consequently, the effects chain 5004 can be controlled by a user-desired addition or modification of one or more effects, by the system applying equipment to a basic track, or a combination thereof.
After the distortion effects and / or equipment selection are applied, the track is preferably transmitted to the equalizer module (5006), which applies a set of equalizer settings to the track. The
201
<img file="MX345588B_D0155.tif" />
track is then preferably transmitted to the .......... module
compression (5008) where a set of compression effects are applied. The EQ and compression settings that
<img file="MX345588B_D0156.tif" />
<img file="MX345588B_D0157.tif" />
will be applied are preferably pre-configured for each musician avatar, although they can also be set or adjusted manually. By applying the above effects, the musical track can be processed in order to be representative of the style, sound and musical tendencies of the virtual musician and the instrument selected by the user.
Once the musician role effects have been applied, a series of producer role effects are applied, as illustrated in Figures 51 and 52. Referring first to Figure 51, a track (5102) is divided between three parallel signal paths, with a separate level control (5104a-c) being applied to each path. It is desirable to have isolated level controls for each route because each route can have different dynamics. Applying the effects in parallel minimizes composition and inappropriate or undesirable effects in the chain. For instruments such as drums, which can include a kick drum, snare drum, hi-hat, cymbal, etc., the audio associated with each drum, hi-hat, cymbal, etc. can be considered as a separate track, where each of those tracks is divided into three Texas routes *<sup>1</sup> for processing.
As shown in Figure 51, a separate effect is then applied to each of the three signal paths. The first path faith is provided to the utility effects module (5106), which applies one or more utility settings to the track. Examples of useful settings include but are not limited to effects such as equalization settings and compression settings. The second route is sent to a delay effects module 5108 that applies one or more delay settings to the track in order to shift the timing of various notes. The third route is sent to a reverb effect module (5110) that applies a set of reverb settings to the track. Although not illustrated, multiple reverb or delay settings can also be applied. The settings for each utility, delay, and reverb effects are preferably pre-configured for each virtual producer selectable through the game interface, although they can also be manually adjustable. Once the utility, delay, and reverb effects are applied, the three path signals are mixed back into a single path by the mixer (5112).
As shown in figure 52, the tracks for each instrument in a single composition
203
<img file="MX345588B_D0158.tif" />
music files are fed to mixer 5202 where they are mixed into a single compilation track. In this way the user can configure the relative volume of the various components (for example, instruments) and can be
<img file="MX345588B_D0159.tif" />
<img file="MX345588B_D0160.tif" />
adjusted to each other to highlight one instrument over another. Each producer can also be associated with unique mix settings. For example, a hip hop style producer may be associated with mix settings that result in louder bass, while a rock producer may be associated with mix settings that result in louder guitars. Once mixed, the compilation track is sent to the equalizer module (5204), the compression module (5206), the limiter module (4708) where the equalization settings, compression settings, limiter settings, respectively, are applied to the compilation track. These settings are preferably pre-configured for each virtual producer selectable by a user selectable user avatar, although they can also be adjusted manually.
In one embodiment, each virtual musician and producer can also be assigned an influence value indicative of their ability to influence a musical composition. These values can be used to determine how the effects described above are applied.
For example, the stronger the influence value of the musician or producer, the greater the impact their settings will have on the music. A similar scenario can also be applied for the role of the producer. For effects that are applied in both the roles of the musician and the producer, such as EQ and compression settings, the influence value can also be used to determine how to reconcile differences in effect settings. For example, in one mode a weighted average of the effect settings can be applied based on the differences in the influence values. As an example, suppose that the influence value can be a number from 1 to 10. If a selected musician has an influence value of 10 and is working with a producer with an influence value of 1, all the effects associated with the selected musician can be applied in full. If the selected musician has an influence value of 5 and you are working with a producer with an influence value of 5, the effect of any of the musician's settings can be combined with the producer's settings in a way that can be random, but it would be preferable if it is predetermined. If the selected musician has an influence value of 1, only a very minimal effect can be applied. In another mode, the associated effect settings can be selected solely based on
<img file="MX345588B_D0161.tif" />
<img file="MX345588B_D0162.tif" />
<img file="MX345588B_D0163.tif" />
which of the virtual musician or producer has a higher influence value.
The effects described in Figures 49-52 can also be applied to any device in the system. For example, in a client-server configuration as described above, the effect settings can be processed either on the server or on the client. In one embodiment, the identification of where to process the bills can be determined dynamically based on the capabilities of the client. For example, if the client is determined to be a smartphone, most of the items can be preferentially processed on the server, while if the client is a desktop computer, most of the items should be processed preferentially on the client .
The foregoing description and drawings merely explain and illustrate the invention, and the invention is not limited thereto. While the description is described in relation to certain implementation or modalities, many details are set forth for the purpose of illustration. Thus, the foregoing simply illustrates the principles of the invention. For example, the invention can have other specific forms without departing from its essential spirit or characteristic. The provisions described are illustrative and not restrictive. For those skilled in the art, the invention is susceptible to additional implementations or modalities and
206 some of these details described in this application can vary considerably without departing from the basic principles of the invention. Therefore, it will be appreciated that those skilled in the art are capable of elaborating various arrangements which, although not explicitly described or shown herein, incorporate the principles of the invention and are therefore within its scope and scope. spirit.
· '4' V '^ 3;
<img file="MX345588B_D0164.tif" />
207
Contents38
220 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88 Sheet 89 Sheet 90 Sheet 91 Sheet 92 Sheet 93 Sheet 94 Sheet 95 Sheet 96 Sheet 97 Sheet 98 Sheet 99 Sheet 100 Sheet 101 Sheet 102 Sheet 103 Sheet 104 Sheet 105 Sheet 106 Sheet 107 Sheet 108 Sheet 109 Sheet 110 Sheet 111 Sheet 112 Sheet 113 Sheet 114 Sheet 115 Sheet 116 Sheet 117 Sheet 118 Sheet 119 Sheet 120 Sheet 121 Sheet 122 Sheet 123 Sheet 124 Sheet 125 Sheet 126 Sheet 127 Sheet 128 Sheet 129 Sheet 130 Sheet 131 Sheet 132 Sheet 133 Sheet 134 Sheet 135 Sheet 136 Sheet 137 Sheet 138 Sheet 139 Sheet 140 Sheet 141 Sheet 142 Sheet 143 Sheet 144 Sheet 145 Sheet 146 Sheet 147 Sheet 148 Sheet 149 Sheet 150 Sheet 151 Sheet 152 Sheet 153 Sheet 154 Sheet 155 Sheet 156 Sheet 157 Sheet 158 Sheet 159 Sheet 160 Sheet 161 Sheet 162 Sheet 163 Sheet 164 Sheet 165 Sheet 166 Sheet 167 Sheet 168 Sheet 169 Sheet 170 Sheet 171 Sheet 172 Sheet 173 Sheet 174 Sheet 175 Sheet 176 Sheet 177 Sheet 178 Sheet 179 Sheet 180 Sheet 181 Sheet 182 Sheet 183 Sheet 184 Sheet 185 Sheet 186 Sheet 187 Sheet 188 Sheet 189 Sheet 190 Sheet 191 Sheet 192 Sheet 193 Sheet 194 Sheet 195 Sheet 196 Sheet 197 Sheet 198 Sheet 199 Sheet 200 Sheet 201 Sheet 202 Sheet 203 Sheet 204 Sheet 205 Sheet 206 Sheet 207 Sheet 208 Sheet 209 Sheet 210 Sheet 211 Sheet 212 Sheet 213 Sheet 214 Sheet 215 Sheet 216 Sheet 217 Sheet 218 Sheet 219 Sheet 220
72 members in 11 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 13194806 | United States of America | – | |
| 201113194806 | United States of America | A | |
| 2012048883 | United States of America | W | |
| 13194806 | – | – | – |
| PCTUS2012048883 | – | – | – |
| US201113194806 | – | – | – |
| WO2012US48883 | – | – | – |
Members72
| Document | Office | Kind | |
|---|---|---|---|
| US2010305732A1 | United States of America | A1 | |
| CA2764042A1 | Canada | A1 | |
| CA2996784A1 | Canada | A1 | |
| US2010307321A1 | United States of America | A1 | |
| WO2010141504A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2010319517A1 | United States of America | A1 | |
| US2010322042A1 | United States of America | A1 | |
| US2011065748A1 | United States of America | A1 | |
| CA2774357A1 | Canada | A1 | |
| WO2011034915A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2438589A1 | European Patent Office (EPO) | A1 | |
| AU2010295698A1 | Australia | A1 | |
| MX2011012749A | Mexico | A | |
| CN102576524A | China | A | |
| EP2477623A1 | European Patent Office (EPO) | A1 | |
| US2012297958A1 | United States of America | A1 | |
| US2012297959A1 | United States of America | A1 | |
| US8338686B2 | United States of America | B2 | |
| US2013025437A1 | United States of America | A1 | |
| JP2013505247A | Japan | A | |
| CA2843437A1 | Canada | A1 | |
| WO2013028315A1 | World Intellectual Property Organization (WIPO) | A1 | |
| CA2843438A1 | Canada | A1 | |
| WO2013039610A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US8492634B2 | United States of America | B2 | |
| US2013220102A1 | United States of America | A1 | |
| US2014053710A1 | United States of America | A1 | |
| US2014053711A1 | United States of America | A1 | |
| US2014140536A1 | United States of America | A1 | |
| EP2737474A1 | European Patent Office (EPO) | A1 | |
| EP2737475A1 | European Patent Office (EPO) | A1 | |
| US8779268B2 | United States of America | B2 | |
| US8785760B2 | United States of America | B2 | |
| CN103959372A | China | A | |
| CN104040618A | China | A | |
| MX2014001192A | Mexico | A | |
| MX2014001194A | Mexico | A | |
| IN741CHN2014A | India | A | |
| IN743CHN2014A | India | A | |
| EP2737474A4 | European Patent Office (EPO) | A4 | |
| CA2929213A1 | Canada | A1 | |
| WO2015066204A1 | World Intellectual Property Organization (WIPO) | A1 | |
| HK1200588A1 | Hong Kong, China | A1 | |
| HK1201975A1 | Hong Kong, China | A1 | |
| EP2737475A4 | European Patent Office (EPO) | A4 | |
| US9177540B2 | United States of America | B2 | |
| US9251776B2 | United States of America | B2 | |
| US9257053B2 | United States of America | B2 | |
| US9263021B2 | United States of America | B2 | |
| US9293127B2 | United States of America | B2 | |
| US9310959B2 | United States of America | B2 | |
| EP2438589A4 | European Patent Office (EPO) | A4 | |
| EP3059886A1 | European Patent Office (EPO) | A1 | |
| EP3063618A1 | European Patent Office (EPO) | A1 | |
| CN106023969A | China | A | |
| CN104040618B | China | B | |
| CN106233245A | China | A | |
| EP2737475B1 | European Patent Office (EPO) | B1 | |
| MX345588BThis record | Mexico | B | |
| MX345589B | Mexico | B | |
| BR112014002269A2 | Brazil | A2 | |
| BR112014002270A2 | Brazil | A2 | |
| MX2016005646A | Mexico | A | |
| EP3063618A4 | European Patent Office (EPO) | A4 | |
| CN103959372B | China | B | |
| EP3059886B1 | European Patent Office (EPO) | B1 | |
| CA2764042C | Canada | C | |
| BRPI1014092A2 | Brazil | A2 | |
| CA2929213C | Canada | C | |
| CN106233245B | China | B | |
| BR112012006027A2 | Brazil | A2 | |
| CN106023969B | China | B |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG | |
| Correction or change in generalHH | HH |
Numbers
- Publication
- 345588
- Publication, DOCDB
- 345588
- Publication, EPODOC
- MX345588
- Application
- 2014001194
- Application, DOCDB
- 2014001194
- Application, EPODOC
- MX202014001194
Titles2
- Spanish
- SISTEMA Y MÉTODO PARA PROVEER AUDIO PARA UNA NOTA REQUERIDA UTILIZANDO UNA MEMORIA CACHÉ DE PROCESAMIENTO.
- English
- SYSTEM AND METHOD FOR PROVIDING AUDIO FOR A REQUIRED NOTE USING A PROCESSING CACHE MEMORY.
Classification
- CPC, 10
- G09B15/00
- G06F3/0481
- G09B15/04
- G10H1/0025
- G10H2210/125
- G10H2220/111
- G10H2220/126
- G10H2230/015
- G10H2230/021
- G10H2240/075
- IPC, 4
- G06F3 0481
- A63H5 00
- G01H7 00
- G09B15 00