Method for search in an audio database
Abstract
This record has no abstract on file.
Term
Term ended
Expired 26 July 2021, 5.2 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
28 claims: 25 independent, 3 dependent
- 1オーディオ・サンプルを識別する方法であって、 前記オーディオ・サンプルの内容に基づき計算される当該オーディオ・サンプルにおける時間上の特定の位置で生じるランドマークと、当該オーディオ・サンプルの前記特定の位置またはその付近における1または2以上の特徴量を含むフィンガープリントとで形成される、サンプル・ランドマーク/フィンガープリント・ペアを生成するステップと、 1または2以上のオーディオ・ファイルの各々に対し、当該オーディオ・ファイルの内容に基づき計算される当該オーディオ・ファイルにおける時間上の特定の位置で生じるランドマークと、当該オーディオ・ファイルの前記特定の位置またはその付近における1または2以上の特徴量を含むフィンガープリントとで形成される、ファイル・ランドマーク/フィンガープリント・ペアを生成するファイル・フィンガープリント生成ステップと、 各サンプル・ランドマーク/フィンガープリント・ペアと過去に生成されたファイル・ランドマーク/フィンガープリント・ペアとのほぼリニアな対応関係を特定する特定ステップと、 前記ほぼリニアな対応関係が多数あるときにベストなファイルを識別する識別ステップと、 を有することを特徴とする方法。
- 2各フィンガープリントは、各ランドマーク位置またはその位置からわずかにオフセットした位置における当該オーディオの多数の特徴量を表現することを特徴とする請求項 1に 記載の方法。
- 3各フィンガープリントは、前記サンプルの時間伸縮に対し不変であることを特徴とする請求項1 又は2 に記載の方法。
- 4各フィンガープリントは、スペクトル・スライス・フィンガープリント、マルチスライス・フィンガープリント、LPC係数、ケプストラム係数、およびスペクトルピークの周波数成分のうちのいずれかとして計算されることを特徴とする請求項1 乃至3 のいずれか 1項 に記載の方法。
- 5前記スペクトル・スライス・フィンガープリントは、ランドマーク時点に対する時間オフセットのセットにおいて計算されることを特徴とする請求項 4 に記載の方法。
- 6各ランドマークの位置は、サウンド記録における特徴的かつ再現可能な位置をみつけるランドマーク分析方法を用いて識別されることを特徴とする請求項1 乃至5 のいずれか 1項 に記載の方法。
- 7前記ランドマーク分析方法は、スペクトルLpノルムを用いて前記サウンド記録においてとりうるすべての時点で瞬時パワーを計算し、ランドマークとしての極大点を選択することを特徴とする請求項 6 に記載の方法。
- 81または2以上のランドマークは、固定もしくは可変のオフセットでの複数の時間スライスにわたるスペクトル成分から得られるマルチスライス・ランドマークであることを特徴とする請求項 6 または 7 に記載の方法。
- 9前記ファイル・ランドマーク/フィンガープリント・ペアはデータベースに格納され、そのデータベース内の各ファイルはそのファイルのフィンガープリントによってインデックスされることを特徴とする請求項1 乃至8 のいずれか 1項 に記載の方法。
- 10前記インデックスはフィンガープリントに基づきソートされることを特徴とする請求項 9 に記載の方法。
- 11固有のフィンガープリントとそれに対応するランドマークのリストへのポインタと含むエントリを有するマスター・インデックス・リストが構成されることを特徴とする請求項1 0 に記載の方法。
- 12各ファイルはSOUND_IDによって識別され、前記データベースは、フィンガープリントと、ランドマークと、SOUND_IDとの3つ組を複数記憶することを特徴とする請求項 9乃至11 のいずれか 1項 に記載の方法。
- 13統計的に最もリニアな対応関係にあるファイル・ランドマーク/フィンガープリント・ペアを有するファイルが前記ベストなファイルとして選択されることを特徴とする請求項1 乃至12 のいずれか 1項 に記載の方法。
- 14サンプル・ランドマーク/フィンガープリント・ペアが許容範囲内でファイル・ランドマーク/フィンガープリント・ペアにマッチするときに、サンプル・ランドマークとファイル・ランドマークのペア(landmark n , landmark* n )においてリニアな対応関係が生じることを特徴とする請求項1 乃至13 のいずれか 1項 に記載の方法。
- 15個々のフィンガープリント同士がマッチし、なおかつ、個々のランドマーク同士がリニアな関係にあるときに、サンプル・ランドマーク/フィンガープリント・ペアとファイル・ランドマーク/フィンガープリント・ペアとの間にリニアな対応関係が生じることを特徴とする請求項1 乃至14 のいずれか 1項 に記載の方法。
- 16フィンガープリント同士が同一または所定の許容範囲内の差にあるときに、両者のフィンガープリントがマッチしたとすることを特徴とする請求項1 5 に記載の方法。
- 17リスト内のサンプル・ランドマークとファイル・ランドマークのペア(landmark n , landmark* n )が landmark* n =m*landmark n +offset、 の関係にあるときに、リニアな対応関係が生じることを特徴とする請求項1 5 または1 6 に記載の方法。
- 18前記サンプルは、音響波、無線波、ディジタルオーディオPCMストリーム、圧縮ディジタルオーディオストリーム、インターネットストリーミング放送のいずれかの形式であることを特徴とする請求項1 乃至17 のいずれか 1項 に記載の方法。
- 19前記サンプル・フィンガープリントはローリング・バッファに格納されることを特徴とする請求項1 乃至18 のいずれか 1項 に記載の方法。
- 20前記特定ステップおよび識別ステップは、前記ローリング・バッファの内容に対して周期的に実行されることを特徴とする請求項 19 に記載の方法。
- 21前記特定ステップおよび識別ステップは、認識するのに十分な情報が前記ローリング・バッファから得られしだい実行されることを特徴とする請求項 19 または2 0 に記載の方法。
- 22前記特定ステップおよび識別ステップは、まずファイルの第1のサブセットに対して実行され、その第1のサブセットでベストなファイルが識別されなかったときに、残りのファイルを収めた第2のサブセットが検索されることを特徴とする請求項1 乃至21 のいずれか 1項 に記載の方法。
- 23前記第1のサブセットは、当該第1のサブセットに含まれないファイルよりも識別される経験的に確率が高いファイルを含むことを特徴とする請求項2 2 に記載の方法。
- 24前記特定ステップは、前記対応する位置間の差分をとり、その差分のヒストグラムのピークを計算することにより、前記対応する位置の散布図における斜線を特定することを特徴とする請求項1に記載の方法。
- 25前記識別ステップは、多数の対応関係を生じる前記ベストなファイルにおける位置に対するオフセットの指標を提供するステップを更に有することを特徴とする請求項1に記載の方法。
- 26オーディオ・サンプルを識別する方法であって、 クライアントからの要求に応じて、前記オーディオ・サンプルの少なくとも一部を、請求項1に記載の各ステップを実行するサーバに取り次ぐステップと、 前記サーバがベストなファイルを識別したことに応じて、前記クライアントに応答を返すステップと を有することを特徴とする方法。
- 27請求項1 乃至26 のいずれか 1項 に記載の方法をコンピュータに実行させるためのプログラムを格納したコンピュータ読み取り可能な記憶媒体。
- 28請求項1 乃至25 のいずれか 1項 に記載の方法を実行するコンピュータシステムであって、 キャプチャされた信号サンプルのランドマーク/フィンガープリント・ペアを含む特徴抽出サマリを、認識処理を実行するサーバ端末に送信するクライアント端末を含むことを特徴とするコンピュータシステム。
Independent claims28
1 paragraph, as filed
[0001] (Field of invention) The present invention relates to information retrieval based on content, and specifically to recognition of an acoustic signal including a sound or a musical sound that is greatly distorted or mixed with high-level noise. [0002] (Background technology) There is a growing need for automatic recognition of music or other audio signals generated from various sound sources. For example, owners and advertisers of works with copyrighted works are interested in obtaining data on the broadcast frequency of their works. Music tracking services provide playlists for major radio stations in the large market. Consumers want to identify songs and promotions that will be aired on the radio so that they can purchase new songs, songs of interest, or other products or services. Both types of continuous acoustic recognition or on-demand acoustic recognition are inefficient and require a large amount of labor when performed manually. Therefore, automatic recognition of music and sound will bring tremendous benefits to consumers, artists, and various industries. As the form of music sales shifts from store purchases to downloads over the Internet, it becomes a real reality to directly link to music recognition performed on computers by Internet purchases or other Internet-based services. Traditionally, the recognition of a song aired on the radio is a playlist provided by either a radio station or a third-party source that matches the time with the radio station where the song was played. This method is essentially limited to radio stations that can receive information. Another option is to rely on embedding an inaudible code in the broadcast signal. The embedded signal is decoded by the receiver and the identification information about the broadcast signal is extracted. The disadvantages of this method are that it requires a dedicated decoding device to identify the signal and that it can only identify songs with embedded chords. [0003] Large-scale audio recognition requires some sort of content-based audio retrieval that identifies identical or similar database signals by comparing unidentified broadcast signals with a database of known signals. Note that content-based audio search differs from audio search by existing web search engines, where only the metadata text that surrounds or accompanies the audio file is searched. Voice recognition is a voice signal It should also be noted that it is useful for indexing and converting into searchable text by well-known techniques, and that many audio signals, including music and sound, cannot be applied. [0004] In a sense, audio information retrieval is similar to text-based information retrieval provided by search engines. However, audio recognition is not similar in that audio signals lack easy-to-identify entities such as words that provide identifiers for retrieval and indexing. For this reason, current audio retrieval methods index audio signals by calculating perceptible features that indicate the quality and characteristics of various types of signals. [0005] Content-based audio retrieval generally involves analyzing a query signal to obtain a large number of representative features and then applying the obtained features to a similarity measurement to locate a database file that most closely resembles the query signal. It is done in. The similarity of the received objects is necessary to reflect the selected perceptual features. Many content-based search methods have been proposed in the past. For example, US Pat. No. 5,210,820 by Kenyon discloses a signal recognition method in which a received signal is processed and sampled to obtain a signal value at each sampling point. Then, the statistical moment of the sample value is calculated in order to generate a feature vector comparable to the identifier of the stored signal and obtain a similar signal. U.S. Pat. Nos. 4,450,531 and 4,843,562 by Kenyon et al. Disclose a similar method of classifying broadcast information that calculates the cross-correlation between unidentified signals and stored reference signals. [0006] JT is a system that searches audio documents by acoustic similarity. "Content-Based Retrieval of Music and Audio," by Foote (C. -CJ Kuo et al., Editor, Multimedia Storage and Archiving Systems II, Proc. of SPIE, volume 3229, pages 138-147, It is disclosed in 1997). A feature vector is calculated by parameterizing each audio file to the Mel scale cepstrum coefficient, and a quantized tree grows from the parameterized data. To execute the query, the unknown signal is parameterized to obtain a feature vector sorted to the leaf node of the quantized tree. A histogram is obtained for each leaf node, resulting in an N-dimensional vector representing an unknown signal. The distance between these two vectors indicates the similarity between the two sound files. In this method, the controlled quantization scheme learns the classification of acoustic features based on classes to which training data are manually assigned, ignoring insignificant fluctuations. Depending on this classification system, different acoustic features are selected as important. Therefore, this method is more suitable for calculating the similarity between songs and classifying music than for recognizing music. [0007] A content-based analysis, storage, retrieval, and segmentation method for audio information is disclosed in US Pat. No. 5,918,223 by Blum et al. In this method, a large number of acoustic features of each file such as loudness, bass, pitch, brightness, bandwidth, and mel frequency cepstrum coefficient are measured periodically. Once the statistical measurements of these features are obtained, they are combined to form a feature vector. Audio data files in the database are searched based on the similarity between their feature vector and the feature vector of the unidentified file. [0008] The serious problem with all the conventional audio recognition methods mentioned above is If the signal you are trying to recognize is subject to linear or non-linear distortion caused by background noise, transmission errors and defects, interference, band-limited filtering, quantization, time warping, voice quality digital compression, etc., it tends to fail. There is. In the conventional method, when the distorted acoustic sample is processed to obtain the acoustic features, the features extracted for the original recording are only a small part. Therefore, the obtained feature vector is not very similar to the feature vector of the original recording, and it is unlikely that the recognition will be performed accurately. Therefore, there is still a need for an acoustic recognition system that works well in the presence of high levels of noise and distortion. [0009] Another problem with conventional methods is that the amount of computation is large and cannot be reduced successfully. Therefore, real-time recognition is not possible with the conventional method using a large-scale database. In such a system, it is not practical to have a database of hundreds or thousands of recordings. Search times in traditional methods tend to extend linearly with the size of the database, making it economically infeasible to extend thousands of recordings to millions. Kenyon's method also requires large-scale dedicated digital signal processing hardware. [0010] Most of the current commercialized methods require strict requirements on the input sample so that they can be recognized. For example, the entire song, or at least 30 seconds of the song, is required to be sampled, or it is required to be sampled from the beginning of the song. It is also difficult to recognize a mixture of multiple songs in one stream. All of these drawbacks hinder the practical application of conventional methods. [0011] (Purpose and effect of the invention) Therefore, a main object of the present invention is to provide a method of recognizing a high level noise or distorted audio signal. [0012] Another object of the present invention is to provide a real-time feasible recognition method that requires only a few seconds for a signal to be identified. [0013] Another object of the present invention is to provide a recognition method capable of recognizing sound based on a sample at an arbitrary position rather than from the beginning of sound. [0014] It is further an object of the present invention to provide a recognition method that does not require encoding an acoustic sample or associating it with a particular radio station or playlist. [0015] Furthermore, it is an object of the present invention to provide a recognition method capable of recognizing each of a mixture of a plurality of sound recordings in one stream. [0016] Furthermore, it is also an object of the present invention to provide an acoustic recognition system in which unknown acoustics can be provided to the system from any environment by almost any known method. [0017] (Overview) These objectives are achieved by recognizing media samples, such as audio samples, from a database index of a large number of known media files. The database index represents the characteristics of the indexed media file at a particular location. A media file is the media file (best) if the relative position of the fingerprint of the media file in the database matches the relative position of the fingerprint of the unknown media sample most faithfully. Media file) is recognized. For audio files, the time progression of the best file fingerprint matches the time progression of the sample fingerprint. [0018] This method is preferably implemented by a distributed computer system and has the following steps: That is, the step of calculating a set of sample fingerprints at a particular location in the sample, [0019] The acquisition step of obtaining a set of file fingerprints that characterizes at least one file position in the media file, the step of identifying a matching fingerprint from the database index, and the specific position and file of the sample. It has a step of generating a correspondence between the file positions and the corresponding positions having an equivalent fingerprint, and a step of identifying the media file when many of the correspondences are in a linear relationship. .. If many of the correspondences are linear, it is considered the best media file. One way to identify a file from many correspondences is to look at the diagonal lines in the scatter plot generated from the correspondence pairs. In one embodiment, the identification of media files from a linear correspondence check is performed from only the first subset of media files. Files in the first subset are more likely to be identified than files not included in the first subset. The probability of identification preferably reflects not only the frequency of prior identification, but also the newness of empirical frequency or timing in past identification. If the media file is not identified from the first subset, the second subset containing the remaining files is searched. Alternatively, the files may be ranked by probability and searched in that rank order. The search ends when the file is identified. [0020] Preferably, the particular position of the sample is reproducibly calculated depending on the sample. Such reproducibly calculated positions are called "landmarks". The fingerprint is preferably numerical. In one embodiment, each fingerprint represents many features of the media sample at each position or at a position slightly offset from that position. [0021] [0021] This method is particularly useful for recognizing an audio sample, in which case the particular location is at the time of the audio sample. At these points in time, for example, the spectral Lp norm of the audio sample is maximized. Fingerprints can be calculated by analysis of various audio samples, preferably fingerprints are invariant to time expansion and contraction of the samples. Examples of fingerprints include spectral slice fingerprints, multi-slice fingerprints, LPC coefficients, cepstrum coefficients, and frequency components of spectral peaks. [0022] The present invention also includes a landmark analysis object that calculates a specific position, a fingerprint analysis object that calculates a fingerprint, a database index that includes the file position and fingerprint of a media file, and an analysis object. Provide a system to realize. The parse object identifies the matching fingerprint in the database index, generates a correspondence, analyzes the correspondence, and selects the best media file. [0023] Also provided is a program storage device accessible by the computer that tangibly embodies a program of instructions that can be executed by the computer to perform each step of the method. [0024] In addition, the invention also provides a method that includes the following steps to index a large number of audio files in a database. That is, it has a step of calculating a set of fingerprints at a specific position of each file and a step of storing the fingerprint, position and ID of the file in memory. Corresponding fingerprints, locations, and IDs are associated in triplicate format. Preferably, the current position of the audio file is calculated and reproducible based on the file. For example, at that point, the spectral Lp norm of the audio file is maximal. In some cases, each fingerprint is preferably a number and represents a number of features of the file near a particular location. Fingerprints can be calculated from any analysis of audio files and digital signal processing. Examples of fingerprints include spectral slice fingerprints, multi-slice fingerprints, LPC coefficients, cepstrum coefficients, frequency components of spectral peaks, and concatenated spectral peaks. [0025] The present invention also provides a method of identifying audio samples that incorporates time-invariant fingerprints and various hierarchical searches. [0026] (Detailed explanation) The present invention provides a method of recognizing a foreign media sample given to a database containing a large number of known media files. Also provided is a method of generating a database index that enables efficient retrieval using the recognition method of the present invention. Although the following description primarily refers to audio data, the methods of the invention include any type of media, including text, audio, video, images, various multimedia combinations of different media types, and the like. -It will be understood that it can also be applied to samples and media files. For audio, the present invention includes samples containing high levels of linear or non-linear distortion caused by, for example, background noise, transmission errors and omissions, interference, band-limited filtering, quantization, time warping, voice quality digital compression, etc. It is especially useful for recognizing. As will be clear from the following description, the present invention can accurately recognize a distorted signal even if only a small part of the calculated features can withstand the distortion. Works with. Any type of audio, including sound, voice, music, or a combination thereof, can be recognized by the present invention. Examples of audio samples include recorded music, radio broadcast programs, and advertisements. [0027] Foreign media samples, as used herein, refer to segments of media data of any size obtained from various sources, as described below. For recognition to take place, the sample must be a rendition of a portion of the index media file in the database used by the present invention. The index media file is the original recording, and the sample can be thought of as a distorted and / or shortened variant or reproduction of the original recording. In general, the sample corresponds only to a small portion of the index file. For example, recognition can occur in 10 seconds of a 5-minute song indexed in the database. The term "file" is used for indexed building blocks, which may be in any format as long as the required values (discussed below) are available. Moreover, there is no need to store or access the file after the value is obtained. [0028] FIG. 1 is a diagram conceptually showing the entire steps of the method 10 of the present invention. Each step will be described in detail below. This method identifies the media file in which the relative position of the unique fingerprint is closest to the relative position of a similar fingerprint in the foreign sample as the best media file. After the outpatient sample is captured in step 12, landmarks and fingerprints are calculated in step 14. Landmarks occur at specific locations (eg, point in time) in the sample. The position of the landmark in the sample is preferably determined and reproducible by the sample itself (ie, depending on the same quality). This means that the same landmark will be calculated at each time even if the processing is repeated for the same signal. For each landmark, a fingerprint is obtained that characterizes one or more characteristics of the sample at or near that landmark. The degree of approximation of the property to the landmark is defined by the fingerprint method used. In some cases, if a property clearly matches a landmark and does not match a previous landmark and a subsequent landmark, then the property is considered close to that landmark. In another case, the property corresponds to multiple landmarks in close proximity. For example, a text fingerprint can be a word string, an audio fingerprint can be a spectral component, and an image fingerprint can be a pixel RGB value. In the following, two general embodiments of step 14 will be described. One is that landmarks and fingerprints are calculated sequentially, and the other is that landmarks and fingerprints are calculated at the same time. [0029] In step 16, sample fingerprints are used to retrieve a set of matching fingerprints from database index 18. Here, matching fingerprints are associated with landmarks and identifiers in a set of media files. The set of identifiers and landmark values for the retrieved file will generate a corresponding pair containing the sample landmarks (calculated in step 14) at the location where the same fingerprint was calculated and the landmarks for the retrieved file. Used in (step 20). The obtained corresponding pairs are sorted by music ID to generate a set of correspondences between sample landmarks and file landmarks of each applicable file. Each set is inspected for alignment of file landmarks with sample landmarks. This means that a linear correspondence is confirmed between the landmarks, and the set is scored according to the number of pairs in the linear correspondence. A linear correspondence occurs when a linear equation can be described in which the correspondence between a large number of sample positions and the file positions is approximately equal within the margin of error. For example, if the gradients of many equations describing a set of corresponding pairs fall within the range of + -5%, then all sets of correspondence are considered to have a linear relationship. Of course, another suitable margin of error may be selected. The ID of the set with the highest score (ie, the highest number of linear correspondences) is identified and returned as the best file ID in step 22. [0030] Further, as will be described later, recognition can be performed by a time component proportional to the logarithm of the number of entries in the database. Even a large database can basically be recognized in real time. That is, it can be recognized with a small time delay after the sample is obtained. This method can identify acoustics based on segments of 5-10 seconds, or even shorter 1-3 seconds. In a preferred embodiment, the landmark and fingerprint analysis of step 14 is performed at the same time that the sample is captured in step 12. The database search (step 16) is performed as soon as the sample fingerprint becomes available, the results of the correspondences are calculated, and the correspondences are periodically inspected for a linear correspondence. Therefore, all steps of this method can be performed simultaneously rather than sequentially as shown in FIG. Note that this method is partly similar to a text search engine. The user presents a search sample and returns a match of the files indexed in the acoustic database. [0031] The method is typically implemented as software that operates on a computer system, with individual steps being most efficiently implemented as independent software modules. Therefore, the system that realizes the present invention consists of landmark and fingerprint analysis objects, indexed databases, and analysis objects for searching database indexes, calculating correspondences, and identifying the best files. Can be done. When the landmark analysis and the fingerprint analysis are performed sequentially, the landmark analysis object and the fingerprint analysis object may be separate objects. Computer instructions for the individual objects are stored in the memory of one or more computers and executed by one or more computer processors. In one embodiment, the code objects are housed together in a single computer system, such as an Intel-based personal computer or other workstation. In a preferred embodiment, the method is implemented by a central processing unit (CPU) distributed by a network, where separate software objects are executed by separate processors to distribute the computational load. Alternatively, each CPU may hold a copy of all software objects, taking into account the isomorphic network of elements, each configured in the same way. In this latter configuration, each CPU has a subset of the database index and is capable of responding to searches for a subset of its own media files. [0032] FIG. 2 is a diagram showing an example of a preferred embodiment of the distributed computer system 30. However, the present invention is not limited to a specific hardware system. System 30 includes a multiprocessing bus architecture 34 or a network protocol such as the Beowulf cluster computing protocol, or Linux-based processors 32a-32f connected by a combination of the two. In such a configuration, the database index is preferably stored in random access memory (RAM) of at least one node 32a in the cluster to ensure a fast start-up of fingerprint analysis. Computation nodes for other objects, such as the marking nodes 32c and 32f, the fingerprinting nodes 32b and 32e, and the alignment check node 32d, require as much RAM as the node 32a that holds the database index. is not it. The number of compute nodes assigned to each object can be increased or decreased as required so that one object does not become a bottleneck. Therefore, the computing network can be highly parallelized, and in addition, multiple signal recognition searches distributed among the available computing resources can be processed simultaneously. This means that it enables applications that return results in a short time even if a large number of users request recognition. [0033] In another embodiment, some functional objects are more tightly coupled, while not so tightly coupled with other objects. For example, landmark analysis objects and fingerprint analysis objects can be placed physically separated from the rest of the objects. One example of this is the strong linkage between landmark and fingerprint analysis objects and signal capture processing. In this deployment configuration, landmark analysis objects and fingerprint analysis objects can be used, for example, on mobile phones, wireless application protocol (WAP) browsers, personal information terminals (PDAs), other remote terminals such as audio search engine client terminals, and so on. It can be embedded as additional hardware or software that is embedded. In Internet-based audio search services such as content display services, landmark analysis objects and fingerprint analysis objects are linked software instruction sets or software, such as Microsoft® Dynamic Link Library (DLL). -It can be incorporated into the client's browser application as a plug-in module. In these embodiments, the objects including all of signal capture, marking, and fingerprinting constitute the client terminal of the service. The client terminal sends a summary of the captured signal sample from which the feature quantity including the landmark and the fingerprint pair is extracted to the server terminal that executes the recognition. Sending a summary of this feature to the server instead of the raw capture signal is advantageous in that the amount of data can be significantly reduced, often to one-500th or less. Such information can be reared on a narrow bandwidth unilateral channel, for example, not with or with the audio stream sent to the server. Can be sent at Le Time. This makes it possible to realize the present invention on a public communication network in which each user is given a narrow bandwidth. [0034] This method will be described in detail below with reference to audio samples and audio files indexed in the acoustic database. This method is roughly divided into two types: construction of an acoustic database index and sample recognition. [0035] (Building a database index) As a prerequisite for acoustic recognition, it is necessary to build a searchable acoustic database index. As used herein, a database refers to a collection of indexed data, but is not limited to a database on a commercial basis. In a database index, elements of connected data are related to each other, and individual elements can be used to retrieve related data. The acoustic database index contains an index set of each file or recording in a selected set or library, including audio, music, advertising, sonar signals, or other acoustics. Each record also has a unique identifier, sound_ID. The sound database itself does not need to store the audio file for each recording, but the sound_ID can be used to retrieve the audio file from another location. Acoustic database indexes are expected to be very large in size as they contain indexes for millions and possibly billions of files. New records should be gradually added to the database index. [0036] FIG. 3 shows a preferred method 40 for constructing a searchable acoustic database index according to the first embodiment. In this embodiment, the landmark is calculated first, and then the fingerprint at or near the landmark position is calculated. It will be apparent to those skilled in the art that other methods of constructing database indexes are possible. In particular, many of the steps listed below are selective, but serve to generate database indexes that can be searched more efficiently. Search efficiency is important for real-time acoustic recognition from a large database, but a small database can be searched at a relatively high speed even if it is not optimally sorted. [0037] To index the acoustic database, each record in the set is landmark-analyzed and fingerprint-analyzed to generate an index set for each audio file. FIG. 4 is a diagram schematically showing a segment of a sound recording in which landmarks (LM) and fingerprints (FP) are calculated. Landmarks occur at specific points in the sound and have a value of time offset from the beginning of the file, and fingerprints characterize the sound at or near a particular landmark location. Thus, in this embodiment, each landmark for a particular file is unique, while the fingerprint can occur multiple times in one file or multiple files. [0038] In step 42, each sound recording is landmark analyzed using a method of finding characteristic and reproducible positions in the sound recording. Suitable landmark analysis algorithms can mark the same point in time in a sound recording in the presence of noise and other linear and non-linear distortions. Some landmark analysis methods are conceptually independent of the fingerprinting process described below, but can be selected to optimize the performance of that process. By landmark analysis, the point in time in sound recording {landmark<sub>k</sub>} Is obtained, and the fingerprint is calculated sequentially at each point in time. According to a good landmark analysis scheme, approximately 5-10 marks are made per second of sound recording. Of course, the density of landmarks depends on their activity in sound recording. [0039] Various methods can be used to calculate landmarks, all of which are within the scope of the present invention. Since the specific method used to realize the landmark analysis scheme of the present invention is known, detailed description thereof will be omitted. A simple landmark analysis method called the power norm calculates the instantaneous power at each point in the sound recording and selects the maximum point. One way to do this is to calculate the envelope by directly adjusting and filtering the waveform. Another method is to calculate the Hilbert transform (orthogonal transform) of the signal and use the Hilbert transform and the squared amplitude of the original signal. [0040] The power norm method of landmark analysis is suitable for finding transitions in acoustic signals. The power norm is a very special case (p = 2) of the general spectral Lp norm. The general spectrum Lp norm is calculated at each point in the acoustic signal, for example by computing a short spectrum by the Fast Fourier Transform (FFT) of the Hanning window. A preferred embodiment uses a sampling rate of 8000 Hz, an FFT frame size of 1024 samples, and a shift width of 64 samples for each time slice. The Lp norm of each time slice is calculated as the sum of the absolute values of the spectral components to the p-th power, and then the p-th root is also calculated selectively. As mentioned above, the landmark is selected as the maximum value obtained at each time. An example of the spectral Lp norm method is shown in FIG. FIG. 5 is a graph of the L4 norm as a time function of an acoustic signal. The dashed line at the maximum value indicates the position where the landmark is selected. [0041] When P = , the L norm is effectively the maximum norm. That is, the norm value is the absolute value of the largest spectral component on the spectral plane. This norm provides a robust landmark and good recognition performance, and is also suitable for tonal music. [0042] Alternatively, a "multi-slice" spectral landmark can be calculated by summing the absolute values of the spectral components to the p-th power over multiple time slices at fixed or variable offsets rather than a single time slice. By finding the maximum value of this extended sum, it becomes possible to optimize the arrangement of the multi-slice fingerprint described later. [0043] Once the landmarks have been calculated, in step 44 the fingerprints are calculated at each landmark time position. A fingerprint is generally a value or set of values that aggregates a set of features at or near each landmark time position in a sound recording. In this preferred embodiment, each fingerprint is a single number that is a hash function of multiple features. Possible types of fingerprints include spectral slice fingerprints, multi-slice fingerprints, LPC coefficients, and cepstrum coefficients. Of course, any type with a fingerprint that characterizes a signal or a signal near a landmark is within the scope of the present invention. Fingerprints can be calculated by any type of digital signal processing or frequency analysis of the signal. [0044] Frequency analysis is performed near each landmark time position to extract the vertices of several spectral peaks to generate a spectral slice fingerprint. A simple fingerprint value is only the frequency value of the largest spectral peak. When such a simple peak is used, very good recognition performance can be obtained even in the presence of noise. However, single frequency spectrum slice fingerprints tend to produce false positives more than other fingerprint analysis schemes because they are not unique. It is possible to reduce the number of false positives by using a fingerprint consisting of a function of 2 or 3 maximum spectral peaks. However, they can be very vulnerable to noise if the second largest spectral peak is not large enough to be discernible for the noise present. That is, the calculated fingerprint value may not be robust enough for reliable reproduction. Nevertheless, the performance of this case is also good. [0045] To take advantage of many acoustic time progressions, a set of time slices is calculated by adding a set of time offsets to the landmark points. For each time slice obtained, a spectral slice fingerprint is calculated. The resulting set of fingerprint information is then combined to form a one-slice or multi-slice fingerprint. Since each multi-slice fingerprint follows a time progression, it is more unique than a single spectral fingerprint, reducing matching failures in database index searches as described below. Experiments have shown that due to the increased eigencity, multi-slice fingerprints calculated from a single maximum spectral peak in each of the two time slices speeded up subsequent database index search calculations (approximately 100). (Double speed), but the recognition rate decreased in the presence of large noise. [0046] It is also possible to use a variable offset to calculate the multi-slice fingerprint instead of using a fixed offset from a given time slice. The variable offset to the selected slice is the offset to the next landmark for the fingerprint, or the offset to the landmark in a range of offsets from the "anchor" landmark. In this case, the time difference between the landmarks is also encoded in the fingerprint along with the multi-frequency information. By adding dimensions by fingerprinting, the degree of uniqueness is further increased and the chance of matching failure is reduced. [0047] In addition to the spectral components, other spectral characteristics may be extracted and used as a fingerprint. Linear Predictive Coding (LPC) analysis extracts predictable linear properties of a signal, such as spectral peaks in addition to spectral shapes. LPC is well known in the field of digital signal processing. For the present invention, the LPC coefficient of the waveform slice fixed at the landmark position can be used as a fingerprint by hashing the quantized LPC coefficient to the index value. [0048] The cepstrum coefficient is useful as a measure of periodicity and can be used to characterize signals with a tonal structure such as speech and many musical instruments. Cepstrum analysis is well known in the field of digital signal processing. For the present invention, a large number of cepstrum coefficients are hashed to the index and used as a fingerprint. [0049] FIG. 6 shows another embodiment 50 in which landmarks and fingerprints are calculated simultaneously. Steps 42 and 44 in Figure 3 have been replaced by steps 52, 54 and 56. As described below, in step 52, a multidimensional function is calculated from the sound recording and landmarks (54) and fingerprints (56) are extracted from the function. [0050] In the embodiment of FIG. 6, landmarks and fingerprints are calculated from the spectrogram of audio recording. A spectrogram is a time-frequency analysis of speech recording, typically using the Fast Fourier Transform (FFT) to spectrally analyze windowed and overlapping frames of acoustic samples. As mentioned earlier, a preferred embodiment uses a sampling rate of 8000 Hz, an FFT frame size of 1024 samples, and a shift width of 64 samples for each time slice. An example of the spectrogram is shown in Figure 7A. The horizontal axis is time and the vertical axis is frequency. Each consecutive FFT frame is arranged vertically at equal intervals along the time axis. This spectrogram shows the energy density at each time-frequency point. The black part on the graph shows high energy density. Spectrograms are well known in the field of audio signal processing. For the present invention, landmarks and fingerprints are obtained from convex angle points by points such as the maxima of the spectrum marked with a circle in the spectrogram of FIG. 7B. For example, the time-frequency coordinates for each peak are obtained, the time is taken as a landmark, and the frequency is used in the calculation of the corresponding fingerprint. The spectral peak landmark approximates the L norm, where the maximum absolute value of the norm determines the landmark position. However, in the spectrogram, the search for maxima is dominated by patches in the time-frequency plane, rather than over the entire time slice. [0051] In the following, the set of convex angle points obtained from the convex angle point extraction analysis of sound recording is referred to as constellation. For constellations with maxima, a suitable analysis is to select the point with the highest energy in the time-frequency plane over the vicinity of each selected point. For example, (t<sub>0</sub>-T, f<sub>0</sub>-F), (t<sub>0</sub>-T, f<sub>0</sub>+ F), (t<sub>0</sub>+ T, f<sub>0</sub>-F), (t<sub>0</sub>+ T, f<sub>0</sub>In a quadrangle with + F) as the apex, that is, a quadrangle with side lengths 2T and 2F by T and F chosen to obtain an appropriate number of constellations, the coordinates (t).<sub>0</sub>, f<sub>0</sub>When the point of) is the point of maximum energy, that point is selected. The square range can also be resized according to the frequency value. Of course, any shape of the region may be used. The maximum energy reference can also be weighted so that the peaks of the competing time-frequency energies are weighted conversely according to the distance in the time-frequency plane, i.e., the distant points are less weighted. For example, energy is weighted by the following equation. [0052] S (t, f) / (1 + C<sub>t</sub>(tt<sub>0</sub>)<sup>2</sup>+ C<sub>f</sub>(ff<sub>0</sub>)<sup>2</sup>) [0053] However, S (t, f) is the square amplitude value at the point (t, f), C<sub>t</sub>And C<sub>f</sub>Is a positive value (it does not have to be a constant). Other distance weighting functions may be used. Other (non-maximum) convex point feature extraction methods may be applied to the constraints of maximal point selection, which are also within the scope of the present invention. [0054] This method yields a set of values that closely resembles a single frequency spectrum fingerprint with many of the same properties described above. The spectrogram time-frequency method produces more landmark / fingerprint pairs than the single frequency method, but it can also cause many matching failures in the matching stage described below. However, it provides more robust landmark and fingerprint analysis than single frequency spectrum fingerprints. This is because the major noise contained in the acoustic sample does not spread to every part of the spectrum in each slice. That is, many landmark-fingerprint pairs in the spectral portion are largely unaffected by major noise. [0055] This spectrogram landmark analysis and fingerprint analysis method is a special case of an analysis method that calculates a multidimensional function of an acoustic signal whose time is one of the dimensions and positions the convex angle point at the value of that function. .. The convex angle point can be a feature value such as a maximum value, a minimum value, or a zero intersection. The landmark is acquired as the time coordinate of the convex angle point and the corresponding fingerprint can be calculated from at least one of the remaining coordinates. For example, the non-time coordinates of the multidimensional convex angle points can be hashed together to form a multidimensional function fingerprint. [0056] The variable offset method for multi-slice spectral fingerprints described above can be applied to spectrograms or other multidimensional function fingerprints. In this case, as shown in the spectrogram of FIG. 7C, the constellation points are connected together to form a link point. Each point in the constellation acts as an anchor point that defines the landmark time, and the remaining coordinate values of the other points are concatenated to form a linked fingerprint. For example, points that are close to each other, as defined later, are connected together to form a more complex overall feature fingerprint that is easier to identify and retrieve. As with multi-slice spectral fingerprints, the purpose of combining information from multiple link convex points into a single fingerprint is to allow the fingerprint values to take more values, thereby matching. To reduce the chance of failure, that is, to reduce the chance that the same fingerprint points to two different music samples. [0057] In principle, each of the N convex points is connected to each other by a two-point connection method that produces approximately N / 2 links. Similarly, for K-point concatenation, the number of possible combinations obtained from the constellation is on the order of N. In order to avoid the expansion of such a combination, it is desirable to limit the neighboring points connected to each other. One way to achieve such a limitation is to define a "target zone" for each anchor point. The anchor point is then connected to a point in its target zone. It is possible to select a subset of points in the target zone so that they are not linked to all points. For example, it is possible to connect only the points related to the maximum peak in the target zone. The target zone may have a fixed shape or a shape that changes according to the characteristics of the anchor point. Anchor point for the constellation of spectral peaks (t)<sub>0</sub>, f<sub>0</sub>A simple example of the target zone in) is that t is the range [t<sub>0</sub>+ L, t<sub>0</sub>+ L + W] is a set of spectral plane points (t, f) such that L is the lead time length and W is the width of the target zone. In this method, all frequencies can be taken within the target zone. For example, L or W can be variable if the rate control mechanism is used to modulate the number of concatenated combinations produced. Instead, for example, the frequency f is in the range [f<sub>0</sub>-F, f<sub>0</sub>Frequency limitation can be achieved by target zone constraints such as within + F] (where F is the boundary parameter). The advantages of frequency limitation are: In psychoacoustics, it is known that melodies tend to interfere well when the sequence of notes has frequencies close to each other. Such constraints make it possible to exhibit more "psychoacoustic-practical" cognitive performance, even though psychoacoustic modeling is not required for the goals of the present invention. f to [f<sub>0</sub>-F, f<sub>0</sub>It is also possible to think of the opposite rule, such as selecting outside the range of [+ F]. This concatenates points of different frequencies so that the constellation extraction results can avoid the case of stuttering time-frequency points that are close in time and have the same frequency. Like other locality parameters, F does not have to be a constant, for example f<sub>0</sub>It may be a function of. [0058] [0058] When the time coordinates do not include the anchor convex point of the fingerprint value, it is necessary to use the relative time value to make the fingerprint time-invariant. For example, the fingerprint can be a function of (i) non-time coordinate values and / or (ii) the difference between the corresponding time coordinate values of the convex angle points. The time difference can be obtained, for example, with respect to the anchor points, or as a sequential difference between a series of convex points in a concatenated set. Coordinates and diffs can be packed into continuous bitfields to form a hashed fingerprint. It will be apparent to those skilled in the art that there are many ways to map a set of coordinate values to fingerprint values and they fall within the scope of the present invention. [0059] A concrete example of this method is the coordinates (t).<sub>k</sub>, f<sub>k</sub>) (However, k = 1, ..., N), N> 1 concatenated spectral peaks are used. And (i) the time of the first peak t<sub>1</sub>Is obtained as the landmark time, (ii) frequency f<sub>k</sub>However, the time difference Δt of the connected peaks at (k = 1, ..., N)<sub>k</sub>= t<sub>k</sub>-t<sub>1</sub>(However, k = 2, ..., N) are both hashed to form a fingerprint value. Fingerprint is Δt<sub>k</sub>-f<sub>k</sub>It can be calculated from all possible points or subsets of coordinates. For example, if desired, all or all of the time difference can be excluded. [0060] Another advantage of forming a fingerprint with multiple points is that the fingerprint coding is invariant, for example, for time expansion and contraction when the sound recording is played at a speed different from the original recording speed. It is in the point that it can be done. This advantage applies to both the spectral method and the time slicing method. Note that in an extended time signal, the time difference and frequency have a contradictory relationship (for example, if the time difference between two points is reduced by half, the frequency doubles). This method uses data from the combination of time difference and frequency in a way that excludes time expansion and contraction from the fingerprint. [0061] For example, the coordinate value (t)<sub>k</sub>, f<sub>k</sub>) (However, in the case of the N-point spectral peak for k = 1, ..., N), the possible intermediate value for hashing to the fingerprint is Δt.<sub>k</sub>= t<sub>k</sub>-t<sub>1</sub>(However, k = 2, ..., N, and fk, k = 1, ..., N). Use any frequency as the reference frequency f<sub>1</sub>By calculating (i) the quotient with other frequencies and (ii) the product with the time difference, the median value can be made invariant with respect to time expansion and contraction. For example, the median is g<sub>k</sub>= f<sub>k</sub>/ f<sub>1</sub>(However, k = 2, ..., N), s<sub>k</sub>= Δt<sub>k</sub>/ f<sub>1</sub>(However, k = 2, ..., N). When the sample is accelerated by α times, the frequency f<sub>k</sub>Is αf<sub>k</sub>, Time difference t<sub>k</sub>Is Δt<sub>k</sub>/ α, so g<sub>k</sub>= αf<sub>k</sub>/ αf<sub>1</sub>, S<sub>k</sub>= (Δt<sub>k</sub>/ α) (αf<sub>1</sub>) = Δt<sub>k</sub>f<sub>1</sub>Will be. These new intermediate values are then combined using a function to form a hashed fingerprint value that is independent of time expansion and contraction. For example, g<sub>k</sub>And s<sub>k</sub>The value of can be hashed by packing it in a concatenated bitfield. [0062] Alternatively, instead of the reference frequency, the reference time difference (eg Δt)<sub>2</sub>) May be used. In this case, (i) the quotient Δt with other time difference<sub>k</sub>/ Δt<sub>2</sub>, (Ii) Product with frequency Δt<sub>2</sub>f<sub>k</sub>, A new intermediate value is calculated. In this case, it is equivalent to using the reference frequency. The resulting value is g above<sub>k</sub>And s<sub>k</sub>Because it can be generated from the product and quotient of. The relativity of frequency ratios can be used exactly the same. That is, the sum and difference of the logarithmic values of the original median can be used instead of the product and quotient, respectively. All time-stretch-independent fingerprint values obtained by such mathematical manipulations (commutations, substitutions, permutations) are within the scope of the present invention. Further, a plurality of reference frequencies or reference time differences for relativizing the time difference may be used. Using multiple reference frequencies or reference time differences is equivalent to using a single reference value. g<sub>k</sub>And s<sub>k</sub>This is because the same result can be obtained by the arithmetic operation of. [0063] The explanation is returned to Fig. 3 and Fig. 6. Landmark analysis and fingerprint analysis by any of the methods described above will give an index set of Sound_ID as shown in Figure 8A. A given sound recording index set is a list of paired values (fingerprint, landmark). Each indexed sound record typically has an order of 1,000 (fingerprint,) in its index set. It has a pair of landmarks). In the first embodiment described above, the landmark analysis and fingerprint analysis methods are basically independent and can be treated as separate and interchangeable modules. Depending on the system, signal quality, and type of acoustics recognized, any of many different landmark analysis or fingerprint analysis modules can be used. In fact, since an index set is a simple composite of value pairs, it is possible and often preferred to use multiple landmark and fingerprint analysis techniques at the same time. For example, certain landmark and fingerprint analysis techniques may be good for detecting unique tone patterns, but may not be sufficient for percussion identification, in which case different algorithms with conflicting attributes may be used. It's good. A technique that uses multiple landmark / fingerprint analysis provides a more robust and sufficient range of recognition performance. Different fingerprint analysis techniques may be used together by ensuring some range for several types of fingerprint analysis. For example, in a 32-bit fingerprint value, the first 3 bits may be used to describe 8 fingerprint analysis techniques and the remaining 29 bits may be used for coding. [0064] An index set is generated for each sound recording that is indexed in the acoustic database. Here, the searchable database index is constructed in such a way as to enable high-speed (ie, logarithmic hours) search. This is achieved in step 46 by building a list of triads (fingerprint, landmark, sound_ID) obtained by adding the sound_IDs corresponding to each of the two in each index set. Such triads are collected as a large index list for all sound recordings. Figure 8B shows an example of this. To optimize the subsequent search process, the triplet list is sorted by fingerprint. Fast sorting algorithms are well known and are discussed in detail in DE Knuth, The Art of Computer Programming, Volume 3: Sorting and Searching, Reading, Mass (Addison-Wesley, 1998). This is incorporated herein by this reference. The high-performance sorting algorithm can sort the list by Nlog N times, where N is the number of entries in the list. [0065] Once the index list has been sorted, step 48 sorts each of the unique fingerprints in the list into a new master index list. Figure 8C shows an example of this. Each entry in the master index list contains a fingerprint value and a pointer to a list of (landmark, sound_ID). Depending on the number and characteristics of sound recordings that are indexed, it is possible that a particular fingerprint will appear hundreds of times or more throughout the collection. Reconstructing the index list into a master index list is optional, but it saves memory because each fingerprint value appears only once. The number of entries in the current list is significantly reduced to a list of unique values, which also speeds up subsequent database searches. Alternatively, the master index list may be constructed by inserting each triad into a balance tree (B-tree). As we all know, there are other ways to build a master index list. The master index list is preferably kept in system memory such as DRAM for fast access during signal recognition processing. The master index list may be kept in the memory of a single node in the system as shown in Figure 2. Alternatively, the master index list may be divided into a plurality of lists and distributed among a plurality of arithmetic nodes. The acoustic database index described above is preferably the master index list shown in FIG. 8C. [0066] The acoustic database index is preferably built offline and additionally updated as new acoustics are input to the recognition system. A new fingerprint may be inserted at the appropriate position in the master index list to update the list. If the new sound recording contains fingerprints that already exist, the corresponding pair (landmark, sound_ID) is added to the list that already exists for those fingerprints. [0067] (Recognition system) Using the master index list generated as described above, acoustic recognition is performed on the input acoustic sample. Acoustic samples are typically supplied by users who are interested in identifying the sample. For example, a user may want to listen to a new song on the radio and find out the artist and title of that song. Samples can be produced from any environment. For example, radio broadcasts, discos, pubs, submarines, sound files, streaming audio segments, stereo systems, etc. And these may include background noise, dropouts, or voice. The user can store the audio sample in a storage device such as a response device, a computer file, a tape recorder, a telephone or mobile telephone, or a voice mail system before supplying the audio sample to the recognition system. As an attachment to any analog or digital sound source (stereo system, television, compact disc player, radio broadcast, answering device, telephone, mobile phone, internet streaming broadcast, FTP, email, based on system and user settings Audio samples are supplied to the recognition system of the present invention from computer files, other devices suitable for transmitting these recordings). Depending on the source, the sample may be in the form of acoustic waves, radio waves, digital audio PCM streams, compressed digital audio streams (Dolby Digital, MP3, etc.), Internet streaming broadcasts, and so on. Users interact with the recognition system through standard interfaces such as telephones, mobile phones, web browsers, and email. Samples will be captured by the system and processed in real time or played back for processing from previously captured sounds (eg sound files). Audio sample is micro during capture It is digitally extracted by a sampling device such as a phone and sent to the system. The sample will be further degraded by channel and sound capture device limitations, depending on the capture method. [0068] When the sound signal is converted to digital format, processing for recognition is performed. As building an index set of database files, landmarks and fingerprints are calculated for the sample using the same algorithm used to process the sound recording database. The method works best if the original sound file is heavily distorted and you still get the same or similar set of landmark and fingerprint pairs as you got for the original recording. .. The index set obtained for the sound sample is a pair set (fingerprint, landmark) of the analytical values shown in FIG. 9A. [0069] Given a pair of analytical values for a sound sample, the database index is searched to identify files that are likely to match. The search is performed as follows. Each (fingerprint) in an index set of unknown samples<sub>k</sub>, landmark<sub>k</sub>) Pairs are fingerprints in the master index list<sub>k</sub>It is processed by searching for. Fast search algorithms for ordered lists are well known and are discussed in detail in DE Knuth, The Art of Computer Programming, Volume 3: Sorting and Searching, Reading, Mass (Addison-Wesley, 1998). Fingerprint of master index list<sub>k</sub>If found, it matches (landmark *<sub>j</sub>, sound_ID<sub>j</sub>) The corresponding list of pairs is copied and landmark<sub>k</sub>In addition, (landmark<sub>k</sub>, landmark *<sub>j</sub>, sound_ID<sub>j</sub>) Form a triad of the form. Here, an asterisk (*) indicates a landmark in one of the index files in the database, and a landmark without an asterisk indicates a sample one. In some cases, it is not necessary to match only when both fingerprints are the same, but if the fingerprints are close to each other (for example, the difference falls within a predetermined threshold). It is preferable to determine that there is. Fingerprints that match by the same thing and fingerprints that match by approximation are referred to here as "equivalent". Sound_ID in triad<sub>j</sub>Corresponds to files with landmarks with an asterisk. Therefore, each triad contains two separate landmarks when the equivalent fingerprint was calculated. One is for database indexes and the other is for samples. This process is repeated for all k over the range of the input sample index set. All the resulting triads are collected in a large candidate list as shown in Figure 9B. The candidate list contains the sound_IDs of the sound file by matching the fingerprints, and these sound_IDs are so called because they are candidates for identification for the input sound sample. [0070] After the candidate list is collected, segmentation by sound_ID is performed. A convenient way to do this is to sort the candidate list by sound_ID or insert the candidate list into a balance tree (B-tree). As mentioned earlier, many sorting algorithms are applicable. The result of this process is a list of candidate sound_IDs, each of which is optionally stripped of its sound_ID, as shown in Figure 9C, as a pair of sample landmark time and database file landmark time.<sub>k</sub>, landmark *<sub>j</sub>) Has a spray list. Therefore, each spray list will contain a set of landmarks that are characterized by fingerprints that are equivalent to each other. [0071] The scatter list for each candidate sound_ID is then analyzed to determine if the sound_ID matches the sample. An optional thresholding step may be used to exclude candidates with very small scatter lists. Obviously, a candidate with only one entry in the scatter list, that is, a candidate with only one fingerprint similar to the sample, will not match that sample. Any suitable threshold of 1 or greater may be used. [0072] Once the final number of candidates is determined, the best candidates are identified. If the best candidate cannot be identified by the following algorithm, a recognition failure message is returned. The key point of this matching process is that the time progress in sound matching should follow a linear correspondence, assuming that the time axes of both are fixed. This is true unless one of the sounds is artificially distorted non-linearly or the playback device is defective, such as a tape deck with anomalies that wiggle in speed. Therefore, a successful landmark pair in the dispersal list for a given sound_ID.<sub>n</sub>, landmark *<sub>n</sub>) Should have a linear correspondence of the following equation. [0073] landmark *<sub>n</sub>= m * landmark<sub>n</sub>+ offset [0074] However, m is a value close to 1 in slope. landmark<sub>n</sub>Is the time point in the input sample, landmark *<sub>n</sub>Is the corresponding time point in the sound recording indexed by sound_ID, and offset is the time offset to the sound recording corresponding to the starting point of the input sound sample. Landmark pairs that fit the above equation with specific values m and offset are said to be in a "linear relationship". Obviously, the concept of linear relationships is valid if there are two or more corresponding landmark pairs. Note that this linear relationship has a high probability of identifying normal sound files, excluding landmark pairs outside the critical range. Although it is possible for two separate signals to contain many identical fingerprints, it is unlikely that these fingerprints will have the same relative time progression. The requirement for a linear correspondence is a key feature of the present invention, and the recognition performance can be greatly improved compared to a technique such as simply counting the total number of ordinary features or measuring the similarity of features. .. In fact, due to this aspect of the invention, even if less than 1% of the original recording fingerprints appear in the input sound sample, that is, the sound sample is very short or heavily distorted. Even so, the sound can be recognized. [0075] Therefore, the problem of determining whether or not to match the input sample is narrowed down to the one that corresponds to finding a diagonal line with a slope of about 1 in the scatter plot of the landmark points in the given scatter list. .. Two examples of scatter plots are shown in FIGS. 10A and 10B. The horizontal axis is the landmark of the sound file, and the vertical axis is the landmark of the input sound sample. In FIG. 10A, a diagonal line with a slope of approximately 1 is recognized. This shows that the song certainly matched the sample, that is, the sound file is the best file. The intercept on the horizontal axis shows the offset in the audio file at the beginning of the sample. There are no statistically significant diagonal lines in the scatter plot in Figure 10B. This indicates that the sound file does not match the input sample. [0076] There are many ways to find diagonal lines in scatter plots, all of which are within the scope of the present invention. The term "locating a diagonal line" refers to all methods that correspond to identifying diagonal lines without explicitly generating diagonal lines. The preferred method is m * landmark from both sides of the above equation.<sub>n</sub>Start by subtracting to obtain the following equation. [0077] (landmark *<sub>n</sub>-m * landmark<sub>n</sub>) = offset [0078] Assuming that m is approximately 1 (that is, there is no time expansion and contraction), the following equation is obtained. [0079] (landmark * n-landmarkn) = offset [0080] [0080] The problem in identifying diagonal lines is narrowed down to finding multiple landmark pairs for a given sound_ID that are split by about the same offset value. This can be easily achieved by collecting a histogram of the offset values obtained by subtracting one landmark from the other landmarks. This histogram may be processed using a fast sorting algorithm or by generating a bin entry with a counter and inserting it into a B-tree. The best offset bin in the histogram contains the maximum number of points. In the following, this bin is referred to as the peak of the histogram. The offset should be positive if the input sound signal is sufficiently contained in a normal library sound file, so landmark pairs with negative offsets are excluded. Similarly, offsets beyond the end of the file are excluded. The number of points in the best offset bin of the histogram is shown for each qualifying sound_ID. This number is the score for each sound recording. The sound record in the highest score candidate list is selected as the best. The best sound_ID to signal successful identification is reported to the user as follows: In order to prevent identification failure, the minimum threshold score may be used to gate control the success of the identification process. If the library sound does not score above the threshold, it will not be recognized and the user will be notified. [0081] When the input sound signal contains multiple sounds, each sound can be recognized. In this case, a plurality of winners are specified in the alignment inspection. You don't need to know that the sound signal contains multiple hits. This is because the alignment test will identify two or more sound_IDs with scores that are significantly higher than the remaining scores. Commonly used fingerprint analysis methods show good linear superposition, so individual fingerprints are extracted. For example, spectrogram fingerprint analysis methods show linear superposition. [0082] When a sound sample undergoes time expansion and contraction, the slopes are not equal to 1. Assuming that the slope of the sample undergoing time expansion and contraction is 1 (assuming the fingerprint is invariant with time expansion and contraction), the calculated offset value will not be uniform. One way to deal with this and adapt to modest time expansion and contraction is to increase the size of the offset bins, i.e. to consider the extent to which the offsets are uniform. In general, if the points are not on a straight line, the calculated offset values will be very different, and a slight increase in the size of the offset bin will not produce a large amount of false positives. [0083] There are other ways to find a straight line. For example, the Radon transform described in T. Risse, "Hough Transform for Line Recognition," (Computer Vision and Image Processing, 46, 327-345, 1989), well known in the field of machine vision and graphics research. The Hough transform may be used. In the Hough transform, each point in the scatter plot projects onto a straight line in (slope, offset) space. Therefore, the set of points in the scatter plot is projected onto a straight line in two spaces in the Hough transform. The peaks in the Hough transform correspond to the intersections of the parameter lines. The global peak of such a given scatter plot transform indicates the maximum number of intersecting straight lines in the Hough transform, i.e. the maximum number of colinearity points. In order to allow 5% speed variation, for example, the configuration of the Hough transform may be limited to the region where the slope parameter fluctuates between 0.95 and 1.05, thereby saving the amount of computation. [0084] (Hierarchical search) In addition to the thresholding step of screening out candidates with a very small spray list, more effective improvements can be taken. One remedy is to segment the database index into at least two sections according to the probability of occurrence, and first search only the sound file that has the highest probability of matching the sample. The division can be done at various stages of the process. For example, it is possible to segment the master index list (Figure 8C) into two or more segments when step 16 or step 20 is performed on any one segment. That is, the file corresponding to the matching fingerprint is searched from only a part of the database index, and the scatter list is generated from that part. If the best sound file is not identified, the process is repeated for another database index. In another implementation, all files are retrieved from the database index, but the diagonal check is done separately on different segments. [0085] Using this technique, computationally intensive diagonal line checking is first performed on a small subset of the sound files in the database index. Since the diagonal line inspection has a time component that is almost linear with respect to the number of sound files to be inspected, it is very effective to perform a hierarchical search. For example, if the sound database index contains fingerprints representing 1,000,000 sound files, but only about 1000 files match frequently searched samples, for example, 95% of search queries Suppose you have 1000 files and only 5% of your search queries are for the remaining 999,000 files. Assuming that the calculation cost is linearly dependent on the number of files, the calculation cost is proportional to 95% of the time for 1000 files and 5% of the time for 999,000 files. Then, the average calculation cost is proportional to about 50,900. Therefore, the hierarchical search can reduce the calculation load to nearly 1/20. Of course, the database index can be segmented into two or more levels (eg, a group of new release songs, a group of recently released songs, a group of old and unpopular songs). [0086] As mentioned above, the search is first performed on the first subset, which is a collection of high-probability files for sound files, and only if this first search fails, on the second subset, which contains the remaining files. Will be done. If the number of points in each offset bin does not reach a predetermined threshold, the diagonal line inspection fails. Alternatively, the two searches may be performed in parallel (simultaneously). If the search for the first subset identifies the correct sound file, a signal is sent to end the search for the second subset. If the search for the first subset does not identify the correct sound file, the search for the second subset continues until the best file is identified. These two different implementations are in the trade-off between computational effort and time. The first implementation example has a light amount of computation, but if the first search fails, it causes a slight delay. In contrast, the second implementation is computationally intensive if the best files are in the first subset, but the delay is minimized otherwise. [0087] The purpose of list segmentation is to estimate the probability that a sound file is the target of a search query and limit the search to the files that are most likely to match the query sample. There are many possible ways to assign probabilities to database sounds and sort them, all of which are within the scope of the present invention. Probabilities are preferably assigned based on recency or frequency when identified as the best sound file. The newness of time is a useful measure, especially for popular songs. As new songs are released, musical interests change very rapidly over time. Once the probability score is calculated, the file is assigned a ranking and the list itself is sorted by that ranking. The sorted list is segmented into two or more subsets for search. A small subset contains a given number of files. For example, if the ranking identifies a file in the top 1000 files, for example, that file is placed in a small subset for fast search. Alternatively, the cutoff points for the two subsets may be adjusted dynamically. For example, all files with scores above a given threshold may be placed in the first subset, which causes the number of files in each subset to change frequently. [0088] A special way to calculate the probabilities is to increment the sound file's score by 1 at each time the query sample is identified as matching. To account for the newness of the timing, all of the scores are periodically revised downward so that the new query results in a stronger ranking than the old query. For example, every time there is a query, all scores can be gradually reduced by a fixed factor. As a result, the score will decrease exponentially if it is not updated. Depending on the number of files in the database (which can easily be as high as 1 million), this method requires a large number of score updates for each query, which is an undesirable situation in some cases. Alternatively, the score may be revised downward at relatively infrequent intervals (eg, once a day). The results of infrequent corrections are similar to, but not exactly the same, the results of making corrections with each query. However, the computational load for updating the ranking is very small. [0089] As a variation of the adjustment due to the novelty of this period, the score update that increases exponentially a<sup>t</sup>(However, t is the elapsed time since the last batch update) may be added to the best sound file each time a query is made. Each time a batch update, a<sub>T</sub>(However, T is the total elapsed time since the last batch update), and all scores are revised downward by dividing. In this variant, a is the recency factor, which is greater than 1. [0090] In addition to the ranking process described above, prior knowledge can be introduced to help determine the strength of the listing. For example, a new release will receive more inquiries than an old song. Therefore, the new release may be automatically placed in the first subset of songs that are likely to match the query sample. This can be done independently of the self-ranking algorithm described above. Using the self-ranking feature as well, the new release will be assigned to one of the initial rankings located within the first subset. New releases can be seeded at the top of the list, at the end of a list of songs with a high probability, or somewhere in the middle of the list. For search purposes, the initial position is not an issue as the ranking will converge over time to reflect the true level of interest. [0091] In the alternative embodiment, the search is performed in the order of ranking of newness of time, and ends when the sound_ID score exceeds a predetermined threshold value. This is equivalent to the above method, where each segment contains only one sound_ID. [0092] Experiments have shown that the best sound score is significantly higher than the scores of all other sound files, so you can choose a suitable threshold with a few experiments. One way to achieve this embodiment is to rank all sound_IDs in the database index, according to the new timing, with arbitrary decisions when the scores are the same. Since each ranking of new time is unique, there is a one-to-one mapping between the new time score and sound_ID. Then, when sorting the sound_ID to create a list of candidate sound_IDs and the accompanying scatter list (Figure 9C), the ranking is s. then used instead of sound_ID . Triad (fingerprint, landmark, An index list of sound_ID) may be generated and the ranking numbers may be combined with the index before the index list is sorted into the master index list. Then, ranking is executed for sound_ID. Alternatively, you can use the search and update features to update the ranked sound_ID. When the ranking is updated, the new ranking is assigned to the old ranking and the mapping is kept consistent. [0093] Alternatively, the rankings may be combined in a later process. Once the scatter list is generated, rankings can be associated with each sound_ID. Then, the set is sorted by ranking. In this implementation, only the pointer to the scatter list needs to be modified. There is no need to repeat grouping into the spray list. The advantage of joining in later processing is that it is not necessary to regenerate the entire database index each time the ranking is updated. [0094] It should also be noted that epidemic rankings can themselves be subject to economic value. That is, the ranking reflects the consumer's preference for checking unknown sound samples. In many cases, the desire to purchase a record of a song directs the inquiry. In fact, if the population information about the user is known, it is possible to implement a different ranking method for each of the requested population groups. A user's population group can be obtained from the profile information that the user received when registering for the recognition service. It is also possible to make a dynamic judgment using standard collaborative filtering technology. [0095] In a real-time system, sound is additionally supplied to the recognition system over time, allowing pipeline recognition. In this case, it is possible to process the input data within the segment to additionally update the sample index set. After each update cycle, a new extended index set is used to search the candidate list of sound recordings by the search and inspection steps described above. The database index is searched for fingerprints that match the newly obtained sample fingerprints, and a new triad (landmark) is searched.<sub>k</sub>, landmark *<sub>j</sub>, sound_ID<sub>j</sub>) Is generated. A new pair is added to the scatter list and a histogram is added. The advantage of this approach is that when enough data is collected to accurately identify the sound recording, for example, the number of points in the offset bin of a sound file exceeds a high threshold, or is the second highest. If the score of the sound file is exceeded, the data collection can be interrupted and the result can be notified. [0096] Once the correct sound is identified, the result is communicated to the user or system in an appropriate manner. The results are, for example, computer printing, email, web search results pages, SMS (short messaging) to mobile phones. service) Can be notified by text messages, computer-generated telephone voice messages, posting of results to websites or internet accounts that users can access later. The notified results will identify the sound, such as song name and artist, classical song composer and recording attributes (eg performers, conductors, venues), advertising companies and products, and various other suitable identifiers. Information may be included. In addition, biographical information, surrounding concert information, and other information of interest to fans may be provided, and hyperlinks to such information may be provided. The notified result may include the absolute score of the sound file or the score in comparison with the next highest scored sound file. [0097] One of the useful achievements of the recognition method is that it does not confuse two different performances of the same sound. For example, if the same classical song is played differently, they will not be considered the same, even if a person cannot detect the difference between the two. This is because it is very unlikely that the landmark / fingerprint pairs and their time progression will match in the two performances. In this embodiment, the landmark / fingerprint pairs must be within about 10ms of each other in order for the linear correspondence to be identified. As a result, the automatic recognition of the present invention provides an appropriate performance / soundtrack or artist / label in all cases. [0098] (Realization example) Hereinafter, continuous sliding window audio recognition, which is a preferred embodiment of the present invention, will be described. A microphone or other source is continuously sampled into the buffer to provide a record of the last N seconds of sound. The contents of the sound buffer are periodically analyzed to determine the ID of the sound content. The sound buffer may be of a fixed size or may be increased according to the size at which the sound is sampled. The latter is called the progressively growing segment of the audio sample. A notification is given to indicate that the sound recording has been identified. For example, log files are collected or displayed on a device that shows song information or purchase information such as title, artist, album cover art, and lyrics. To avoid duplication, you will only be notified when the ID of the recognized sound changes, for example, when the jukebox program changes. Such a device can be used to generate a list of music played from any sound stream (radio, internet streaming radio, hidden microphone, phone call, etc.). In addition to the song ID, information such as the recognition time can be logged. If specific information is available (eg from GPS), then this information can be logged. [0099] Each buffer may be identified from the beginning to achieve identification. Alternatively, the sound parameters may be extracted into, for example, a fingerprint or other intermediate feature extraction format and stored in a second buffer. Prior to the second buffer, a new fingerprint may be added along with the old fingerprint that is discarded from the end of the buffer. The advantage of such a rolling buffer method is that it is not necessary to duplicate the same analysis on overlapping segments of the sound sample, thus saving a lot of computation. The identification process runs periodically on the contents of the rolling fingerprint buffer. For small portable devices, the fingerprint stream has a very large amount of data, so fingerprint analysis should be performed on that device and the results should be sent to the recognition server using a relatively low bandwidth data channel. It may be. The rolling fingerprint buffer may be placed on a portable device and sent to the recognition server each time, or the recognition server may be provided with a rolling fingerprint buffer and recognition sessions may be continuously performed on the server. You may be asked. [0100] In such a rolling buffer recognition system, a new sound recording can be recognized as soon as sufficient information is obtained to recognize it. Sufficient information may be less than the buffer length. For example, a characteristic song can be uniquely recognized by playing for 1 second, and even if the buffer has a length of 15 to 30 seconds, the system will recognize it periodically in 1 second, and the song will be recognized immediately. Can be recognized. Conversely, if a non-characteristic song requires a few more seconds of sample to be recognized, the system will have to wait a long time before declaring the song's ID. In this sliding window recognition method, sounds are recognized as soon as they can be identified. [0101] It is important to note the following: Although the present invention has been described on the premise of a full-featured system and method, the configurations of the present invention can be distributed in the form of computer-readable media containing instructions in various formats, and also. Those skilled in the art will appreciate that the present invention applies regardless of the format of the signal recorded on the media actually used for its distribution. Devices accessible by such computers include computer memory (RAM, ROM), floppy disks, CD-ROMs, and transmission media such as digital or analog communication links. [0102] It will be apparent that the embodiments described above can be modified in many ways without departing from the technical scope of the invention. Therefore, the technical scope of the present invention is defined by the claims and the equivalent scope thereof. [Simple explanation of drawings] FIG. 1 is a flowchart showing a method of the present invention for recognizing an acoustic sample. FIG. 2 is a block diagram showing an example of a distributed computer system that realizes the method of FIG. [Fig. 3] It is a flowchart which shows the method of constructing the database index of an acoustic file. FIG. 4 is a diagram schematically showing landmarks and fingerprints calculated for an acoustic sample. FIG. 5 is a graph of the L4 norm of an acoustic sample showing landmark selection. FIG. 6 is a flow chart illustrating another embodiment of constructing a database index of acoustic files used in the method of FIG. [Fig. 7A], [Fig. 7B], FIG. 7C is a spectrogram showing convex points and connected convex points. [Fig. 8A], [Fig. 8B], FIG. 8C shows an index set, an index list, and a master index list according to the method of FIG. [Fig. 9A], [Fig. 9B], FIG. 9C shows an index list, a candidate list, and a scatter list according to the method of FIG. [Fig. 10A], FIG. 10B is a scatter plot showing successful and unsuccessful identification of unknown acoustic samples, respectively.
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP09138691A | Cites | Japan |
| JP62159195A | Cites | Japan |
| US05918223A | Cites | United States of America |
| JP03291752A | Cites | Japan |
| JP62073298A | Cites | Japan |
| JP2001075992A | Cites | Japan |
| US04582181A | Cites | United States of America |
| WO01088900A1 | Cites | World Intellectual Property Organization (WIPO) |
| WO01004870A1 | Cites | World Intellectual Property Organization (WIPO) |
48 members in 14 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 22202300 | United States of America | P | |
| 22202300 | United States of America | P | |
| 60222023 | United States of America | – | |
| 09839476 | United States of America | – | |
| 83947601 | United States of America | A | |
| 83947601 | United States of America | A | |
| 0108709 | European Patent Office (EPO) | W | |
| 0108709 | European Patent Office (EPO) | W | |
| 2000222023 | – | – | – |
| 2001839476 | – | – | – |
| 2001008709 | – | – | – |
| US20000222023P | – | – | – |
| US20010839476 | – | – | – |
| WO2001EP08709 | – | – | – |
Members48
| Document | Office | Kind | |
|---|---|---|---|
| WO0211123A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU8976601A | Australia | A | |
| WO0227600A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU9298201A | Australia | A | |
| WO0211123A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US2002083060A1 | United States of America | A1 | |
| EP1307833A2 | European Patent Office (EPO) | A2 | |
| BR0112901A | Brazil | A | |
| KR20030059085A | Republic of Korea | A | |
| HK1051248A | Hong Kong, China | A | |
| HK1051248A1 | Hong Kong, China | A1 | |
| WO0227600A3 | World Intellectual Property Organization (WIPO) | A3 | |
| JP2004505328A | Japan | A | |
| US2004199387A1 | United States of America | A1 | |
| CN1592906A | China | A | |
| US6990453B2 | United States of America | B2 | |
| EP1307833B1 | European Patent Office (EPO) | B1 | |
| US2006122839A1 | United States of America | A1 | |
| AT329319T | Austria | T | |
| ATE329319T1 | Austria | T1 | |
| DE60120417D1 | Germany | D1 | |
| DK1307833T3 | Denmark | T3 | |
| PT1307833E | Portugal | E | |
| DE60120417T2 | Germany | T2 | |
| ES2266254T3 | Spain | T3 | |
| CN1996307A | China | A | |
| KR100776495B1 | Republic of Korea | B1 | |
| US7346512B2 | United States of America | B2 | |
| US2008208891A1 | United States of America | A1 | |
| CN100538701C | China | C | |
| CN1592906B | China | B | |
| US7853664B1 | United States of America | B1 | |
| US7865368B2 | United States of America | B2 | |
| US2011071838A1 | United States of America | A1 | |
| US8190435B2 | United States of America | B2 | |
| JP4945877B2This record | Japan | B2 | |
| US2012221131A1 | United States of America | A1 | |
| US8386258B2 | United States of America | B2 | |
| US2013138442A1 | United States of America | A1 | |
| US8700407B2 | United States of America | B2 | |
| US8725829B2 | United States of America | B2 | |
| US2014316787A1 | United States of America | A1 | |
| BRPI0112901B1 | Brazil | B1 | |
| US9401154B2 | United States of America | B2 | |
| US2016328473A1 | United States of America | A1 | |
| US9899030B2 | United States of America | B2 | |
| US2018374491A1 | United States of America | A1 | |
| US10497378B2 | United States of America | B2 |
29 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Written request for registration of change of domicileJAPANESE INTERMEDIATE CODE: R313531S531 | S531 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Notification of acceptance of power of attorneyJAPANESE INTERMEDIATE CODE: R3D02RD02 | RD02 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Notification of change in applicantJAPANESE INTERMEDIATE CODE: A711A711 | A711 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| Notification of change in applicantJAPANESE INTERMEDIATE CODE: A711A711 | A711 |
Numbers
- Publication
- 4945877
- Publication, DOCDB
- 4945877
- Publication, EPODOC
- JP4945877B
- Application
- 2002516764
- Application, DOCDB
- 2002516764
- Application, EPODOC
- JP20020516764
Titles2
- Japanese
- 高い雑音、歪み環境下でサウンド・楽音信号を認識するシステムおよび方法
- English
- Systems and methods for recognizing sound / musical signals in high noise and distortion environments
Classification
- CPC, 8
- G06F16/634
- G11B20/10
- G10L19/018
- G10L17/26
- G11B27/28
- G10L15/26
- G06F16/683
- G10L25/54
- IPC, 5
- G10L11 00
- G10L15 00
- G10L15 10
- G06K9 00
- G10L15 26