System and method for automatic audio content analysis for word spotting indexing classification and retrieval
Abstract
According to the present invention, a system for indexing an audio stream for subsequent information retrieval using special audio pre-filtering, skimming the audio stream, gisting and summarizing the audio stream, but indexing only the appropriate speech segments generated by the speech recognition engine, and A method is provided. Special indexing features are described that improve the search and reproducibility of the information retrieval system used after indexing word spotting. The present invention involves translating an audio stream into intervals, each interval comprising one or more segments. For each interval segment, it is determined whether the segment exhibits one or more predetermined audio characteristics, such as a special zero crossing rate range, a specific energy range, and a specific spectral energy concentration range. Audio characteristics are qualitatively determined to represent each audio event including silence, music, speech and lyrics. In addition, interval groups are determined whether or not they match qualitatively predefined meta-patterns such as constant continuous speech, final ideas, stuttering and speech emphasis, etc., and then the audio stream is determined based on interval classification and meta-pattern matching. is indexed, where appropriate features are indexed to improve subsequent retrieval capabilities following information retrieval. In addition, the long-term alternatives generated by the speech recognition engine are indexed with their respective weights to improve subsequent reproducibility.

Term
Term ended
Projected expiry passed 19 January 2020, 6.7 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
23 claims: 4 independent, 19 dependent
- 1오디오 신호를 분석하기 위한 컴퓨터 실시 방법에 있어서, 한 개 이상의 세그먼트의 임시 순서를 각각 포함하는 오디오 신호의 한 개 이상의 간격 내의 오디오 이벤트를 검출하는 단계와, 음성 경계부를 관련된 신뢰 레벨로 식별하기 위해 오디오 이벤트를 분석하는 단계와, 정확도를 개선하기 위해 정성적으로 결정된 규칙을 이용하여 음성 경계부 및 신뢰 레벨에 기초하여 오디오 신호를 색인하는 단계와, 재현 능력을 개선하기 위해 관련된 가중과 함께 오디오 신호 내의 최소한 하나의 워드로 대안을 색인하는 단계와, 이의 색인을 이용하여 오디오 신호를 워드 스폿팅, 요약 및 스키밍하는 것들 중 하나 이상을 보증하는 단계 를 포함하는 컴퓨터 실시 방법.
- 2최소한 하나의 오디오 신호를 분류 및 색인하기 위한 컴퓨터 유용 코드를 가지고 있는 컴퓨터 유용 매체를 포함하는 데이터 기억 장치를 포함하는 컴퓨터에 있어서, 한 개 이상의 세그먼트를 각각 포함하는 간격으로 오디오 신호를 번역하기 위한 논리 수단과, 세그먼트 간격이 최소한 하나의 각각의 오디오 이벤트를 각각 나타내는 한 개 이상의 선정된 오디오 특징을 나타내는 지의 여부를 결정하기 위한 논리 수단과, 결정 수단에 응답하여 간격을 각각의 오디오 이벤트와 관련시킴으로써 간격을 분류하기 위한 논리 수단과, 최소한 하나의 간격 그룹이 미리 정의된 세트의 메타 패턴 내의 메타 패턴과 일치하는 지의 여부를 결정하기 위한 논리 수단과, 간격 그룹이 메타 패턴과 일치하는 것이 결정될 때 간격 그룹을 메타 패턴 분류와 관련시키기 위한 논리 수단과, 간격 분류 및 메타 패턴 분류에 기초하여 오디오 신호를 색인하기 위한 논리 수단 을 포함하는 컴퓨터.
- 3제 2 항에 있어서, 신호로부터의 워드를 표현하기 위해 음성 인식 엔진을 이용하여 오디오 신호의 적절한 부분만을 처리하기 위한 논리 수단과, 최소한 소정의 워드에 대한 대안을 엔진으로부터 수신하기 위한 논리 수단과, 워드 및 대안 중 최소한 소정의 것에 대한 신뢰 레벨을 엔진으로부터 수신하기 위한 논리 수단과, 신뢰 레벨의 최소한의 일부분에 기초하여 워드 및 대안을 색인하기 위한 논리 수단 을 더 포함하는 컴퓨터.
- 4제 3 항에 있어서, 대안이 N개의 문자 보다 길고 x%를 초과하는 신뢰성을 가지고 있는 워드만을 수신하는 컴퓨터.
- 5제 2 항에 있어서, 각각의 선정된 오디오 특징이 최소한의 오디오 신호 부분의 제로 교차율, 최소한의 오디오 신호 부분의 에너지, 최소한의 오디오 신호 부분의 스펙트럼 에너지 집중 및 주파수들 중 한 개 이상에 기초를 두고 있는 컴퓨터.
- 6제 2 항에 있어서, 간격을 분류하기 전에 세그먼트를 정규화하기 위한 논리 수단을 더 포함하는 컴퓨터.
- 7제 6 항에 있어서, 정성적으로 정의된 선정된 세트의 패턴이 일정하게 연결된 음성, 및 음성과 조합된 음악을 포함하는 컴퓨터.
- 8제 6 항에 있어서, 신호의 색인을 이용하여 오디오 신호를 스키밍, 요점 정리 및 요약하기 위한 최소한의 간격 및 메타 패턴 분류 부분을 제공하기 위한 논리 수단을 더 포함하는 컴퓨터.
- 9제 2 항에 있어서, 세그먼트 간격이 한 개 이상의 선정된 오디오 특징을 나타내는 지의 여부를 결정하기 위한 논리 수단이 세그먼트에 관련된 한 개 이상의 오디오 특징이 각각의 임계치와 같은 지의 여부를 간격 내의 각각의 세그먼트에 대해 결정하기 위한 수단과, 각각의 특징이 각각의 임계치와 같을 때 한 개 이상의 오디오 특징과 관련된 한 개 이상의 카운터를 각각 증가시키기 위한 수단과, 한 개 이상의 카운터를 간격 내의 세그먼트의 수와 비교하기 위한 수단, 비교 수단에 기초하여 간격의 분류를 보증하는 간격을 분류하기 위한 논리 수단 을 포함하는 컴퓨터.
- 10제 2 항에 있어서, 미리 정의된 세트의 오디오 이벤트는 색인하기 위한 논리 수단이 이에 기초하여 오디오 신호를 색인할 수 있도록 음성에서의 강조부, 음성에서의 머뭇거림 및 음성에서의 결론 아이디어를 더 포함하는 컴퓨터.
- 11제 10 항에 있어서, 간격을 분류하기 위한 논리 수단에 의해 음성으로서 분류된 최소한 하나의 간격 내의 한 개 이상의 우세 주파수를 결정하기 위한 수과, 한 개 이상의 세그먼트가 우세 주파수의 상한 N%(N은 수이다)를 포함할 때 음성에서의 강조부와 한 개 이상의 세그먼트를 관련시키기 위한 수단과, 한 개 이상의 세그먼트가 우세 주파수의 하한 N%(N은 수이다)를 포함할 때 음성에서의 결론 아이디어와 한 개 이상의 세그먼트를 관련시키기 위한 수단 을 더 포함하는 컴퓨터.
- 12제 11 항에 있어서, 음성에서의 강조부와 모두 관련된 임시 순서 세그먼트가 선정된 기간 보다 긴 기간을 정의하는 지의 여부를 결정하고, 만일 그렇다면 임시 순서 세그먼트를 중요한 아이디어 언어로서 색인하기 위한 수단을 더 포함하는 컴퓨터.
- 13오디오 신호를 분석하기 위한 컴퓨터 실시 방법에 있어서, 한 개 이상의 세그먼트의 임시 순서를 각각 포함하는 오디오 신호의 한 개 이상의 간격 내의 오디오 이벤트를 검출하는 단계와, 오디오 이벤트에 기초하여 오디오 신호를 색인하는 단계와, 이의 색인을 이용하여 오디오 신호를 스키밍, 요점 정리 또는 요약하는 단계 를 포함하는 컴퓨터 실시 방법.
- 14제 13 항에 있어서, 신호로부터의 워드를 표현하기 위해 음성 인식 엔진을 이용하여 오디오 신호의 적절한 부분만을 처리하는 단계와, 최소한 소정의 워드에 대한 대안을 엔진으로부터 수신하는 단계와, 최소한 소정의 워드 및 대안에 대한 신뢰 레벨을 엔진으로부터 수신하는 단계와, 신뢰 레벨에 최소한 일부분 기초하여 워드 및 대안을 색인하는 단계 를 더 포함하는 컴퓨터 실시 방법.
- 15제 13 항에 있어서, 검출 단계가 세그먼트 간격이 최소한 음악 및 음성을 포함하는 최소한 하나의 각각의 오디오 이벤트를 각각 나타내는 한 개 이상의 선정된 오디오 특징을 나타내는 지의 여부를 결정하는 단계와, 결정 수단에 응답하여 간격을 각각의 오디오 이벤트와 관련시킴으로써 간격을 분류하는 단계와, 최소한 하나의 간격 그룹이 미리 정의된 세트의 메타 패턴 내의 메타 패턴과 일치하는 지의 여부를 결정하는 단계와, 간격 그룹이 메타 패턴과 일치하는 것이 결정될 때 간격 그룹을 메타 패턴 분류와 관련시키되, 오디오 신호의 색인이 간격 분류 및 메타 패턴 분류에 기초하여 표현되는 단계 를 포함하는 컴퓨터 실시 방법.
- 16제 15 항에 있어서, 세그먼트 간격이 한 개 이상의 선정된 오디오 특징을 나타내는 지의 여부를 결정하는 단계가 세그먼트에 관련된 한 개 이상의 오디오 특징이 각각의 임계치와 같은 지의 여부를 간격 내의 각각의 세그먼트에 대해 결정하는 단계와, 각각의 특징이 각각의 임계치와 같을 때 한 개 이상의 오디오 특징과 관련된 한 개 이상의 카운터를 각각 증가시키는 단계와, 한 개 이상의 카운터를 간격 내의 세그먼트의 수와 비교하되, 상기 간격 분류 로직 수단이 상기 비교 단계에 기초하여 간격을 분류하는 단계 를 포함하는 컴퓨터 실시 방법.
- 17제 16 항에 있어서, 간격을 분류하는 단계 중에 음성으로서 분류된 최소한 하나의 간격 내의 한 개 이상의 우세 주파수를 결정하는 단계와, 한 개 이상의 세그먼트가 우세 주파수의 상한 N%(N은 수이다)를 포함할 때 음성에서의 강조부와 한 개 이상의 세그먼트를 관련시키는 단계와, 한 개 이상의 세그먼트가 우세 주파수의 하한 N%(N은 수이다)를 포함할 때 음성에서의 결론 아이디어와 한 개 이상의 세그먼트를 관련시키는 단계 를 더 포함하는 컴퓨터 실시 방법.
- 18제 17 항에 있어서, 음성에서의 강조부와 모두 관련된 임시 순서 세그먼트가 선정된 기간 보다 긴 기간을 정의하는 지의 여부를 결정하고, 만일 그렇다면 임시 순서 세그먼트를 중요한 아이디어 언어로서 정의 및 색인하는 단계를 더 포함하는 컴퓨터 실시 방법.
- 19컴퓨터 프로그램 제품에 있어서, 디지털 처리 장치로 판독가능한 컴퓨터 프로그램 기억 장치와, 프로그램 기억 장치 상의 프로그램 수단을 포함하되, 최소한 하나의 오디오 신호를 인덱싱하기 위한 방법을 실행하기 위해 디지털 처리 장치로 수행가능한 명령어를 실시하는 프로그램 코드 요소를 구비하며, 상기 방법 단계는, 한 개 이상의 세그먼트를 각각 포함하는 간격으로 오디오 신호를 번역하는 단계와, 간격 세그먼트가 최소한의 오디오 신호 부분의 제로 교차율, 최소한의 오디오 신호 부분의 에너지, 최소한의 오디오 신호 부분의 주파수 및 최소한의 오디오 신호 부분의 스펙트럼 에너지 집중을 포함하는 한 세트의 특징으로부터 선택되고, 최소한의 음악 및 음성을 포함하는 최소한 하나의 각각의 오디오 이벤트를 각각 나타내는 한 개 이상의 선정된 오디오 특징을 나타내는 지의 여부를 결정하는 단계와, 결정 단계에 응답하여 간격을 각각의 오디오 이벤트와 관련시킴으로써 간격을 분류하는 단계와, 신뢰 레벨에 최소한 일부분 기초하여 워드 및 대안을 색인하는 단계를 포함하는 컴퓨터 프로그램 제품.
- 20제 19 항에 있어서, 상기 방법은 신호로부터의 워드를 표현하기 위해 음성 인식 엔진을 이용하여 오디오 신호의 적절한 부분만을 처리하는 단계와, 최소한 소정의 워드에 대한 대안을 엔진으로부터 수신하는 단계와, 워드 및 대안 중 최소한 소정의 것에 대한 신뢰 레벨을 엔진으로부터 수신하는 단계와, 신뢰 레벨에 최소한 일부분 기초하여 워드 및 대안을 색인하는 단계 를 더 포함하는 컴퓨터 프로그램 제품.
- 21제 19 항에 있어서, 상기 방법은 최소한 하나의 간격 그룹이 미리 정의된 세트의 메타 패턴 내의 메타 패턴과 일치하는 지의 여부를 결정하는 단계와, 간격 그룹이 메타 패턴과 일치하는 것이 결정될 때 간격 그룹을 메타 패턴 분류와 관련시키되, 오디오 신호의 색인이 메타 패턴에 최소한 일부분 기초하는 단계를 포함하는 컴퓨터 프로그램 제품.
- 22제 21 항에 있어서, 상기 방법은 세그먼트에 관련된 한 개 이상의 오디오 특징이 각각의 임계치와 같은 지의 여부를 간격 내의 각각의 세그먼트에 대해 결정하는 단계와, 각각의 특징이 각각의 임계치와 같을 때 한 개 이상의 오디오 특징과 관련된 한 개 이상의 카운터를 각각 증가시키는 단계와, 한 개 이상의 카운터를 간격 내의 세그먼트의 수와 비교하되, 상기 간격 분류 로직 수단이 상기 비교 수단에 기초하여 분류를 행하는 단계를 더 포함하는 컴퓨터 프로그램 제품.
- 23제 22 항에 있어서, 상기 방법은 간격을 분류하는 단계중에 음성으로서 분류된 최소한 하나의 간격 내의 한 개 이상의 우세 주파수를 결정하는 단계와, 한 개 이상의 세그먼트가 우세 주파수의 상한 N%(N은 수이다)를 포함할 때 음성에서의 강조부와 한 개 이상의 세그먼트를 관련시키는 단계와, 한 개 이상의 세그먼트가 우세 주파수의 하한 N%(N은 수이다)를 포함할 때 음성에서의 결론 아이디어와 한 개 이상의 세그먼트를 관련시키는 단계를 더 포함하는 컴퓨터 프로그램 제품.
Independent claims23
28 paragraphs, as filed
A computer implemented method for analyzing an audio signal, and a computer and a computer program product thereof
1 is a schematic diagram showing the present invention;
2 is a flow diagram illustrating the overall indexing logic of the present invention;
3 is a flow diagram illustrating logic for determining an audio characteristic of a segment;
4 is a flow diagram illustrating logic for determining whether a segment is quiet;
5 is a flow diagram illustrating the logic for determining whether a segment is voiced;
Fig. 6 is a flow diagram continuation of the logic shown in Fig. 5;
7 is a flow diagram illustrating the logic for determining whether a segment is music;
Fig. 8 is a flow diagram continuation of the logic shown in Fig. 7;
9 is a flow diagram illustrating the logic for determining whether a segment is lyric;
Fig. 10 is a flow diagram continuation of the logic shown in Fig. 9;
11 is a flow diagram illustrating logic for skimming, recapitulating, and summarizing;
12 is a flow diagram illustrating the logic for another classification and indexing of audio streams based on "events of interest" in words and audio;
13 is a flow diagram illustrating logic for determining whether a speech sample represents an emphasis in speech, a final idea of a speech, or an important idea of a speech;
14 is a flow diagram illustrating the logic for determining whether a harmonic is present;
15 is a flow diagram illustrating a summary resulting from an indexed audio stream;
Fig. 16 is a schematic diagram showing a screen summarizing an audio stream that has been indexed;
Explanation of symbols for the main parts of the drawing
10 : Audio content analysis system 12 : Computer
14 : audio engine 16 : computer diskette
17 : floppy disk drive 18 : video monitor
20 : Printer 22 : Keyboard
24 : Mouse 25 : Data transmission path
26 : Database 28 : Audio Source
29 : speech recognition engine
<background-art><p>BACKGROUND OF THE INVENTION Field of the Invention [0002] The present invention relates generally to audio streams, including audio streams extracted from video, and in particular, classifying audio streams to support subsequent retrieval, gisting, summarizing, skimming, and global searching of audio streams. and systems and methods for indexing.</p><p>With the rapid growth of computer usage in general and multimedia computer applications in particular, large amounts of audio continue to be generated, for example from audio-video applications, and then the audio is stored electronically. As can be seen by the present invention, as the number of audio files increases, it becomes quite difficult to quickly use the stored audio streams and effectively utilize the existing audio file directories or other existing means for accessing them. do. For example, it may be desirable to access an audio stream extracted from a video based on a user query that retrieves information, provides a summary of the audio stream, or allows the user to skim or recapitulate the audio stream. Accordingly, the present invention is based on the recognition that there is an increasing need to effectively search for a particular audio stream that a user requires access to among the thousands of different audio streams that are very preferably stored.</p><p>Conventional information retrieval techniques are based on the assumption that the source text, whether extracted from audio or not, is independent of noise and errors. However, when the source text is extracted from the audio, this assumption becomes meaningless. This is because the speech recognition engine is used to transform the audio stream into computer-remembered text, which presents inaccuracies and inherent impediments to mission performance, so it is important to achieve error-free and noise-free text conversion. It really matters. For example, a given word in an audio stream cannot be recognized accurately at all (saying land can be translated as ramb), diminishing the reproducibility and retrieval capability of the information retrieval system. Since retrieval capability (precision) refers to the ability of a system to retrieve only accurate documents, recall refers to the system's ability to retrieve as many accurate documents as possible. Fortunately, the present invention has recognized that it is possible to illuminate the limiting factors of a speech recognition engine in converting an audio stream into text, and it is possible to improve the retrieval and reproducibility of the information retrieval system by illuminating these limiting factors.</p><p>In addition to the above considerations, the present invention may in various instances allow a user to not only wish to reproduce an audio stream stored in digital form in order to hear the audio stream, but also allows the user to listen to the entire audio stream or only a specific portion thereof. Or, you may want to access information from the audio stream. In practice, the user may only want to hear an audio stream or a summary of the streams, or to understand the gist of the audio stream. For example, the user may want to hear only portions of the audio stream that are done in a particular form or spoken by a particular person, or, in the case of recorded programming, the user may only want to hear portions of the programming that are not commercially available. Likewise, the user may wish to fast forward through the audio. For example, a user may wish to quickly pass a less interesting portion of an audio stream (eg, a commercial) while keeping the portion of interest at an audible rate.</p><p>However, past efforts on audio content analysis, such as those described in Japanese Patent Application Laid-Open Nos. 8063184 and 10049189 and European Patent Application Laid-Open No. 702351, have not only placed great emphasis on the above considerations, but also voice recognition computers. The focus is on improving the accuracy of input devices or simply improving the quality of digitally processed speech. Perhaps in an effort to achieve their intended purpose, their past efforts have paradoxed the inability to consider indexed audio streams based on audio events within the stream to support subsequent navigation, gist and summarization of computer memory audio streams. I couldn't.</p><p>U.S. Patent No. 5,199,077 describes word spotting for sound editing and indexing. This method acts as a keyword index for single-speaker audio or video recordings. The above-mentioned Japanese Patent Laid-Open Nos. 8063184 and 10049189 relate to audio content analysis as a step toward improving speech recognition accuracy. In addition, Japanese Patent Laid-Open No. 8087292A uses audio analysis to improve the speed of a speech recognition system. The above-mentioned European Patent Publication No. EP702351A includes identifying and recording an audio event to help recognize unknown phases and voices. U.S. Patent No. 5,655,058 describes a method for segmenting audio data based on speaker identification, and European Patent Publication No. EP780777A processes audio files with a speech recognition system for extracting spoken words to index audio is described.</p><p>The methods described in these systems are aimed at improving the accuracy and performance of speech recognition. The indexing and retrieval system described is based on speaker recognition, or on the use of words as a search term and direct application to speech recognition on audio tracks. In contrast, the present invention relates to indexing, classifying and summarizing real-world audio, as understood herein, consisting of unambiguous audio consisting of only a single speaker, i.e., speech segments. Recognizing this point, the present invention utilizes all of the systems and methods described below in which music and noise are segmented in speech segments and applied to distinct speech segments achieved with advanced search systems in which speech recognition is made in consideration of audio analysis results. to improve previous word spotting techniques.</p><p>The content of audio including the method described in the document entitled "Content-Based Classification, Search, and Retrival of Audio" (hereinafter referred to as "Muselefish") published in IEEE Multimedia 1996 by Erling et al. Other techniques for analysis have been described. However, the method that Muselefish classifies sounds is not driven by a heuristically determined rule, but by a statistical analysis. As will be appreciated by the present invention, qualitatively determined rules are more powerful than statistical analysis in classifying sounds, and rule-based classification methods can more accurately classify speech than could be statistically based systems. have. Moreover, the Muse Refish system is intended to be used only on short audio streams (less than 15 sec). Therefore, it is not suitable for retrieving information from longer streams.</p><p>Another method for indexing audio (hereinafter referred to as "MoCA") is described, including a method described in a document entitled "Automatic Audio Content Analysis" published in ACM Multimedia 96 (1996) by Pfeiffer et al. has been However, like many similar methods, the MoCA method is a method that seeks to identify audio that is domain specific, ie related to a special type of video event, such as a violence. The present invention recognizes that many audio and multimedia applications benefit from a more generalized ability to segment, classify and search for audio based on its content, particularly one or more selected audio events therein.</p></background-art><tech><p>A method is described for facilitating reliable information retrieval, also referred to as word spots, within a long, unstructured audio stream comprising an audio stream extracted from audio-video data. The present invention provides special audio prefiltering to identify domain/application specific speech boundaries to index only appropriate speech segments generated by speech recognition engines to facilitate reliable subsequent word spotting, among other applications described below. use the To do so, the present invention analyzes the content of the audio stream to identify content-specific, application-specific, type-specific, distinct speech boundaries with the relevant confidence level. Next, the present invention uses the confidence level generated by the speech recognition engine to weight the invention to index transcripts of only selected audio portions (i.e., appropriate speech) as generated by the sound recognition engine. and trust level. Therefore, the present invention not only essentially improves the speech recognition engine, but also improves the search and reproducibility of the information retrieval system (which may use a speech recognition engine) by improving the way audio streams are indexed.</p><p>The present invention provides a visual representation of an audio stream so that the user can browse or skim the stream, play only the audio segments of interest and/or index the audio stream for information retrieval. It may be implemented as a general purpose computer programmed in accordance with the steps of the present invention to classify and index an audio signal, also referred to as an audio stream, comprising audio extracted from a video to subsequently provide a summary to the user.</p><p>The present invention can also be practiced as a manufactured water-machine part that implements a program of instructions that is used by a digital processing device and can be executed by the digital processing device to ensure the current logic. The present invention is embodied in certain mechanical parts that allow a digital processing device to perform the method steps of the present invention. In another aspect, a computer program product readable by a digital processing device and implementing a computer program is described. A computer program product combines a computer readable medium with program code elements that ensure the logic described below. And, computer-implemented methods for executing logic are described herein.</p><p>Thus, in one aspect, a computer-implemented method for analyzing an audio signal comprises detecting audio events in one or more intervals of the audio signal, each interval in any sequence relating to one or more segments. includes The audio event is analyzed to identify the speech boundary with the associated confidence level, and then the method of the present invention indexes the audio signal based on the speech boundary and the confidence level using qualitatively determined rules to improve accuracy. In addition, the method of the present invention uses an index to spot, summarize and skim the audio signal using an alternative, with associated weights, to improve the reproducibility for subsequent guarantees on one word at least one in the audio signal. index by the word of</p><p>In another aspect, a computer for classifying and indexing an audio signal is described. As will be described in detail below, a computer implements computer usable code means comprising logic means for translating an audio signal into intervals, each interval comprising one or more segments. The logic means then determines whether the segment intervals represent one or more predetermined audio features, also referred to as audio features, each audio feature representing at least one respective audio event. Further, the logic means classifies the interval by associating the interval with each audio event in response to the determining means. Furthermore, logical means are provided for determining whether at least one spacing group matches a metapattern within a predefined set of metapatterns, the logical means selecting the spacing group when it is determined that the spacing group matches the metapattern. It relates to meta-pattern classification. At this time, the logic means indexes the audio signal based on the interval classification and the meta-pattern classification.</p><p>In a preferred embodiment, the logic means process only the audio signal part, suitably using a speech recognition engine to represent the words from the signal. The engine generates recognized words and alternatives to them with an associated confidence level. In a simple embodiment, the present invention only indexes longer words (3 characters or more) with a confidence level for recognition of 90% or greater. A more general general-purpose solution is to index recognized words and alternatives based on weights, which weights depend on the level of confidence about the recognition, the confidence level of the alternative word (if any), the length of the recognized word and either of these. change</p><p>Further, in a preferred embodiment, each selected audio characteristic is characterized by a zero crossing rate (ZCR) of the smallest audio signal part, the smallest audio signal part, the spectral energy (SE) distribution and the frequency (F) of the smallest audio signal part. based on one of the Further, in a preferred embodiment, the predefined set of audio events includes music, voice, quiet and lyrics. With respect to metapatterns, the predefined set of patterns includes, but is not limited to, certain unstructured voices (news broadcasts or educational programs) and music combined with voices (commercials), Patterns are defined qualitatively.</p><p>The present invention also envisions classifying and indexing audio streams comprising speech based on "events of interest" in speech, such as the idea of emphasis in sound, hesitation in sound and conclusion ideas in sound. . Accordingly, means are provided for determining a dominant frequency in each sample of a series of samples for at least one interval that is classified as negative. Speech intervals relate to emphasis in the phonetic text when they include the upper limit N% of the dominant frequency, N being a qualitatively determined number, preferably 1. On the other hand, speech intervals are related to the idea of conclusion in speech when they cover the lower bound N% of the dominant frequency. Moreover, if any series of intervals all associated with emphasis in speech define a duration longer than a predetermined duration, then the overall sequence is indexed as an important idea in the sound.</p><p>In a particularly preferred embodiment, logical means are provided for normalizing the segment before classifying the interval. Furthermore, the logic means provide index of intervals and meta-pattern classification for skimming, gisting and summarizing the audio signal using the index of the signal.</p><p>Means are provided for determining whether a segment of an interval exhibits one or more predetermined audio characteristics, whether the one or more audio characteristics associated with the segment equal a respective threshold. If so, the counter related to the audio feature is incremented, and when all segments in the interval are examined, the counter is compared to the number of segments in the interval, and then the interval is sorted based on the comparison.</p><p>In another aspect, the computer program product comprises a computer program storage device readable by a digital processing device, the program means being on the program storage device. The program means to be executed by the digital processing device to perform method steps for indexing the at least one audio signal for a subsequent summary of the signal so that the user can use the summary for browsing and/or playing only the audio type of interest. Contains program code elements that can be According to the present invention, the method steps translate the audio signal into intervals, each interval comprising one or more segments, the segment duration being the zero crossing rate of the smallest audio signal portion, the smallest audio signal portion, the minimum audio signal and determining whether it appears as one or more predetermined audio features selected from a set of features comprising a frequency of the portion and a spectral energy concentration of the minimum audio signal portion. As the present invention contemplates, each audio feature represents at least one respective audio event comprising a minimum of music and voice. The intervals are classified by associating the intervals with the indexed audio signal based at least in part on each audio event and interval classification.</p></tech>
<p>Hereinafter, with reference to the accompanying drawings will be described in detail with respect to embodiments including the advantages, configuration and action of the present invention.</p><p>Referring first to FIG. 1 , a reference numeral 10 is shown for analyzing audio content (including audio temporal image data) to index, classify, and retrieve audio. In the particular architecture shown, system 10 includes a digital processing device, such as computer 12 . In one intended embodiment, computer 12 is a personal computer or computer 12 manufactured by International Business Machines Corporation (IBM) of Armonk, NY, as shown, with an AS400 accompanying an IBM Network Station. It may be any computer including a computer sold under the same brand. Computer 12 may also be a Unix computer or IBM RS/6000 250 workstation with 128 MB of main memory running an OS/2 server or Windows NT server or AIX 3.5.</p><p>Computer 12 includes an audio engine 14 schematically shown in FIG. 1 that can be executed by a processor within computer 12 as a series of computer-executable instructions. Such instructions may reside, for example, in RAM of computer 12 .</p><p>Optionally, the instructions may be included in a data storage device having a computer readable medium, such as the computer diskette 16 shown in FIG. 1 , which may interact with the floppy disk drive 17 of the computer 12 . The instructions may also be stored on a DASD array, magnetic tape, conventional hard disk drive, electronic read-only memory, optical storage device, or other suitable data storage device. In an illustrative embodiment of the invention, the computer executable instructions may be lines of C++ code.</p><p>1 also shows that system 10 includes an output device, such as a video monitor 18 and/or printer 20, and an input device, such as a computer keyboard 22 and/or mouse 24, as is known in the art. may include peripheral computer equipment. Other output devices such as other computers and the like may be used. Similarly, keyboard 22 and mouse 24 and other input devices may utilize, for example, trackballs, keypads, touch screens, and voice recognition devices.</p><p>The computer 12 can access an electronically stored database 26 containing audio data via a data transmission path 25 . Audio data may be input into database 26 from an appropriate audio source 28 . It should be understood that audio data may be input directly into engine 14 from audio source 28, which may be an analog or digital audio source, such as, for example, a broadcast network or radio station. Moreover, the database 26 is stored locally on the computer, in which case the path 25 may be an internal computer bus or the database 26 may be remote from the computer 12, and the path 25 may be local, such as the Internet. It is a communication network or wide area network. For simplicity of explanation, engine 14 accesses speech recognition engine 29 . Speech recognition engine 29 may be any suitable speech recognition engine, such as described, for example, in U.S. Patent No. 5,293,584, assigned to the assignee of the present invention, which is incorporated herein by reference. The speech recognition engine 29 may be the assignee's "Large Vocabulary Continuous Speech Recognition" system of the present invention.</p><p>Illustrative applications of the present invention, namely summary and skimming, may refer to FIG. 15 . Starting at block 300, the received audio stream is indexed using qualitatively defined rules described below. A summary of the indexed audio is then displayed at block 302 whenever requested by the user. This summary 304 appears on the display screen 306 of FIG. 16 , which should be understood that the display screen 306 may be presented on the monitor 18 ( FIG. 1 ). As shown, the summary 304 may include audio types including noise, speech, music, emphasis in speech, laughter, animal (howl), and the like.</p><p>Moving to block 308 of FIG. 15 , a viewing or playback option is selected by the user from the playback options menu 310 ( FIG. 16 ), and the audio selected based on the user selection is reduced, i.e., does not interfere with unselected audio. It is played without receiving. As shown, the user may choose to play the audio of the type selected in block 302 in any order, or may be selected by search capability, ie, reliability or likelihood, in which the audio is actually the selected form. If the user selects a search capability, processing moves to block 312 of FIG. 15 to analyze the indexed audio to play only audio events of interest to the user.</p><p>An identification of the audio being played may be displayed in a playback window 324 on the screen 306 . When audio is extracted from video, the video can be played on window 314 . The user also has a previous button 316 for selecting the previous audio clip, a next button 318 for selecting the next audio clip and a play to listen to the selected clip, i.e. to play the selected clip. button 320 may be selected. However, as described above, the present invention preferably has other applications including information retrieval by word spotting. Irrespective of the application, the ability of the present invention to index audio effectively is to perform with improved reproducibility, making subsequent applications easier to be more accurate in the case of word spotting.</p><p>Accordingly, an audio stream index according to the logic of the present invention will be described below with reference to FIG. 2 . Beginning at block 30 , an audio stream is received by an audio engine 14 . It should be understood that the stream is transformed using a short form Fast Fourier Transform (FFT) function, and then the low amplitude noise component of the FFT is filtered out of the signal before the steps described below.</p><p>Moving to block 31, the stream is divided into arbitrary successive intervals, for example of duration of 2 seconds, each interval being repeatedly divided into one or more segments of duration of 100 milliseconds. However, intervals and segments for different durations may be used within the scope of the present invention.</p><p>From block 31, logic is applied to each segment to determine whether the segment can be optimally classified as one of a predetermined set of audio events by determining the audio characteristics of each segment, as described in more detail below. moves to block 32 for examining . The selected audio event according to a preferred embodiment of the present invention includes silence, voice, and speech on music. If a segment is not classified, it is designated as an unclassified segment.</p><p>The logic proceeds to the next block 33, where each interval is classified by associating the interval with one of the audio events. That is, each interval is correlated to one of the audio events described above based on the result of the inspection of the segment obtained in block 32 . Then, at block 34, whether any interval order (to some extent, sometimes missing intervals) matches one of a set of qualitatively predefined meta-pattern types. is decided The presence of an audio signal or meta-pattern within the stream is identified as an audio stream based on the interval classification obtained in block 33 . For example, a short alternating sequence of 30 seconds of music, voice, and lyrics in a given sequence can match a predefined commercial meta-pattern type, so that a certain qualitatively determined meta-pattern type can be created. It may be classified in the constituting block 33 . Alternatively, the interval classification order of voice-music-music may coincide with a qualitatively predefined meta-pattern to set the education/training type. Other meta-pattern types, such as cartoons and news, can likewise be predefined qualitatively. If necessary, the meta-pattern of the meta-pattern may be predefined qualitatively, such as defining four and only four meta-patterns, which are meta-pattern broadcast news breaks, in turn. Accordingly, a number of meta-pattern types falling within the scope of the present invention can be defined qualitatively. At this time, it can be seen that the meta-pattern is necessarily a predefined order of variously classified intervals.</p><p>From block 35, the process moves to block 36 for processing the selected portion of the audio stream having the speech recognition engine 29 (FIG. 1). The speech recognition engine 29 converts the portion of the audio stream to be processed into text represented by words composed of one or more alpha-numeric characters. Importantly, the entire audio stream required is processed in block 36 . Instead, only portions of the audio stream, eg, classified as news broadcasts at block 35, may be sent to a speech recognition engine for processing. As recognized herein, processing a long unstructured audio stream, which may contain several different types of domain/application speech boundaries having a speech recognition engine, can lead to errors in the output of the speech recognition engine. have. For example, a speech recognition engine may generate a number of errors when trying to convert a segment with speech and music into text. Thus, processing only specific (proper) types of domain/application speech segments reduces errors caused by deficiencies inherent in conventional speech recognition engines.</p><p>Also, as shown in block 36, the selected audio portion is converted to text, but at least some of the words represented, preferably all, are used for two weights called confidence level weights and strong weights. . The weighting is based in part on whether a particular word is extracted from an emphasis segment in speech as described below.</p><p>Then, at block 37, the DO loop must have a minimum length if the word is N characters in the following two situations: preferably N is an integer, e.g. 3, and the word is at least Only words that satisfy the need to return from the speech recognition engine 29 with a confidence level of 90% are provided. The confidence level may be a range of probabilities if desired. Therefore, the present invention utilizes the features of the speech recognition engine to more accurately convert a longer spoken word into text compared to the accuracy of the speech engine when converting a shorter spoken word into text. The step at block 37 may be considered a filter where words of length less than N are not indexed. Alternatively, words of any length may be considered at block 37, with shorter words being removed later and rated relatively low in the search.</p><p>The DO loop proceeds to block 38, where the speech engine 39 interrogates for alternatives to the word during inspection. At block 39, the two alternatives described above are identified as items to be indexed with the word during examination, although all alternatives may be considered in the preferred case. As with the word under examination, weights are assigned alternatively. Similarly, alternative word lattices may be used rather than single word alternatives. Then, at block 40, the stream is indexed using words and alternatives with their respective weights for subsequent retrieval by an information retrieval system, such as, for example, a system known in the art called Okapi. . With the above discussion in mind, it can be seen that only the appropriate speech segments are indexed at block 40 to support subsequent information retrieval of the text based on the question.</p><p>With respect to the search recognized by the present invention, words that do not exist in the vocabulary of the word recognition system are not presented in the generated copy. Therefore, when there is a question, an out-of-vocabulary word cannot return a certain result. With this in mind, a search system such as Okapi may use similar domains (e.g., broadcast news, office access lizards extracted from the corpus of telecommunications and medical).</p><p>As described above, a weight is computed for each word (and alternatives, if any). The weight assigned to a word depends on several factors including the associated confidence level returned by the speech recognition engine, the frequency of the inverse document, and whether the word is strong. In a particularly preferred embodiment, the weight of the word is determined as follows.</p><p>if</p><p>α1 = 0.5 and α2 = 1 + α1 (empirical determination),</p><p>Id = average document length of the lengths of documents d and I,</p><p>qk = kth item according to the question,</p><p>Cd(qk) is the coefficient for question item k in document d,</p><p>ECd(qk) = Edk is the expected coefficient for question item k in document d,</p><p>Cq(qk) = coefficient of kth items in question q,</p><p>Eq(qk) = Eqk expected coefficient of kth item in question q,</p><p>n(qk) = number of documents containing item qk,</p><p>n,(qk) = expected number of documents containing item qk,</p><p>Q, = the total number of items in the question including all alternating words as described above, and N is the total number of documents;</p><p>pi(qk) = weight indicating the confidence level according to the occurrence of ith of the kth question item from the word recognition engine,</p><p>ei(qk) = weight indicating emphasis in speech according to the occurrence of ith in the kth question item</p><p>In case,</p><p>Inverse document frequency for kth question item = idf(qk):</p><p>idf(qk) = log {(Nn,(qk)+α1)/(n,(qk) + α1)} and</p><p>An appropriate score to evaluate document "d" for question "q" = S(d,q):</p><p>S(d,q) = sum of k=1 to Q, of {Edk*Eqk*idf(qk)}/{α1 + α2(Id/I,) + Edk},</p><p>From here,</p><p>Edk = sum of products from I=1 to Q, of {pi(qk)*ei(qk)} for document "d", </p><p>Eqk = sum of products from I=1 to Q, of {pi(qk)*ei(qk)} for question "q".</p><p>When the question is typed and all items have the same emphasis in speech, ei(qk) is a constant, e.g., e. On the other hand, when the user wants to change the emphasis in the voice of an item, the users can type in a prefix symbol such as +word (word). In this case, ei (qk) is It has a default value between 0 and 1 changed to . Since the question is being asked, the logic below to find the emphasis in the voice within the speech is used to determine the strong prefix of each item, given that it has a unique strong item, and ei(qk) takes a value between 0 and 1. Have.</p><p>Figure 3 shows the processing of each segment from the audio stream in more detail. Starting at block 44 , a DO loop is provided in which one or more sound characteristics for each kth segment are determined at block 46 and normalized at block 48 . Specifically, at block 46, the zero crossing rate (ZCR), energy (E) and special energy concentration (RSi) for each segment are determined, as well as the frequency can fall within several predefined ranges i . As discussed below, all or only a subset of these subsets of audio features may be used.</p><p>The zero crossing rate determines the multiple of the segment in which the audio signal amplitude passes the zero value. The energy determines the sum of the squared audio signal amplitudes for each segment. In contrast to this, the particular energy concentration of each segment is established by a plurality of RSi values for each I-th frequency range defined as the sum of the squares of the frequencies in each I-th frequency range provided to the segment. In a preferred embodiment, four frequency ranges are used. By way of example only, the first frequency range R1 is 0-1000 HZ, the second frequency range R2 is 1000-8000 HZ, the third frequency range R3 is 8000-16,000 HZ, and the fourth frequency range is (R4) exceeds 16,000 HZ.</p><p>However, audio features other than the preferred features described above may be used. For example, bandwidth, harmonic (the deviation of the linear spectrum of a sound from the good harmonic spectrum), and luminance (the short form of the Fourier magnitude spectrum stored as logarithmic frequencies) can be tonality. center) can be used.</p><p>At block 48, the computed audio features are satisfactorily normalized. The normalized version of the measured audio feature is the quotient of the difference between the measured audio feature and the mean of that feature across all segments and the standard deviation of that feature for all segments. For example, the normalized spectral energy concentration (NRi) for a segment is given by</p><p>NRi = (RSi - mean(RSi)/SRSi</p><p>Referring now to Figure 4, the logic by which the present invention examines an audio segment is illustrated. Figures 4-10 show a preferred set of qualitative instructional methods with preferred thresholds for specifying various tests for voice, quietness, music, etc., other qualitative instructional methods and/or thresholds being defined. You have to understand that it can be done. Beginning at block 50, a DO loop is provided for each segment within the interval. Proceeding to decision block 52, it is determined whether the frequency percentage of the segment in the first frequency band R1 is greater than 90% compared to all sample frequencies of the segment under examination. When a preferred sampling frequency of 44 KHZ and a segment duration of 100 ms is used, 20 samples per segment can be obtained.</p><p>If more than 90% of the sampled frequencies of the segment are within the first frequency band R1, processing moves to block 54 to name the segment as quiet or otherwise designate or classify it. From block 54 or decision block 52, if the check is negative, logic proceeds to decision block 56 to determine whether the last segment in the interval has been checked, otherwise to get the next segment As logic moves to block 58 , it returns to decision block 52 . However, when the last segment has been checked, the logic ends at state 60 .</p><p>Fig. 5 illustrates a check herein for determining whether a segment is a voiced segment. Beginning at block 62, a DO loop is provided for each segment within the interval. Proceeding to decision block 64, it is determined whether the frequency percentage of the segment in the third frequency band R3 is greater than 15% relative to all sampled frequencies in the segment under examination. In this case, the voice frequency counter is uniformly incremented at block 66 .</p><p>From block 66 or decision block 54, if the test is negative, logic moves to decision block 68 to determine if the zero crossing rate (ZCR) of the segment under test is greater than six. In this case, the negative ZCR counter is uniformly incremented at block 70 . From block 70 and decision block 68, if the check is negative, logic proceeds to decision block 72 to determine whether the last segment in the interval has been checked, otherwise, get the next segment As the logic moves to block 74 for the purpose, it returns to decision block 64 . However, when the last segment has been checked, the logic proceeds to FIG. 6 .</p><p>As can be seen in the present invention, the presence (or absence) of harmonic frequencies in audio can be used to determine whether the audio is music or speech. Typically, spectral analysis is used to determine musical harmonies or segments of chords and some structure of music for note analysis. However, the present invention uses the absence of musical harmonics detected by reliably examining the voice.</p><p>Therefore, as shown in Fig. 6, after checking the segment interval to classify the interval as negative according to the preferred embodiment of the present invention, three situations must be met. In particular, starting at decision block 73, it is determined whether or not the interval is named as a harmonic according to the logic shown in FIG. 14 and described below. If not (if the interval indicates that the interval is voiced), processing moves to decision block 74 where it is determined whether the value of the voice frequency counter exceeds 40% of the number of segments in the interval. In other words, at decision block 74, it is determined whether at least 40% of the segments within the interval under examination satisfy the situation at decision block 64 of FIG. If so, logic moves to decision block 76 to apply a second check of negatives, i.e., to determine whether the value of the negative ZCR counter is less than 20% of the number of segments in the interval under examination. In other words, at decision block 76, it is determined whether less than 20% of the segments within the interval under examination satisfy the situation at decision block 68 of FIG. If either of the checks in decision block 74 of FIG. 6 are not satisfied, or if the interval is found to be harmonic in decision block 73, the logic ends at state 78, otherwise, Intervals are classified as negative in block 80 and indexed before termination. It can now be seen that a confidence level may be generated based on the value of the voice counter at the end of the process of FIG. 6 , which is subsequently used when an interval classified as negative matches the interval sequence with a meta-pattern. It indicates the possibility of voice to do.</p><p>Referring now to FIG. 7 , shown is the check herein for checking whether a segment is music. Beginning at block 82, a DO loop is provided for each segment within the interval. Proceeding to decision block 84, it is determined whether the frequency percentage of the segment in the third frequency band R3 is greater than 15% relative to all sampled frequencies in the segment under examination. If so, the music frequency counter is uniformly incremented in block 86 .</p><p>From either block 86 or decision block 84, if the test is negative, logic moves to decision block 88 to determine whether the zero crossing rate (ZCR) of the segment under test is less than five. If so, the music ZCR counter is uniformly incremented at block 90 . From block 90 or decision block 88, if the test is negative, it is determined whether the normalized third spectral energy concentration NR3 (as determined in block 48 of FIG. 3) of the segment under test exceeds 100,000. Logic proceeds to decision block 92 to determine whether If so, the music spectrum EN counter is uniformly incremented at block 94 . From either block 94 or decision block 92, if the check is negative, logic proceeds to decision block 96 to determine whether the last segment in the interval has been checked; , logic moves to block 98 to return to decision block 84 . However, when the last segment is checked, the logic proceeds to FIG. 8 .</p><p>After checking the segment interval, in order to classify the interval as music, one of three situations must be met. In particular, starting at decision block 100, it is determined whether the value of the music frequency counter exceeds 80% of the number of segments in the interval. If so, logic moves to block 102 to classify the interval as music, index and end the interval. However, if a segment fails the music check in decision block 100, a second check is applied to the music, i.e. whether the value of the music ZCR counter exceeds 95% of the number of segments in the interval being tested. Logic proceeds to decision block 104 to determine If the second check is met, the logic categorizes the interval as music at block 102; otherwise, the logic moves to decision block 106 to apply a third check for the music.</p><p>In decision block 106, it is determined whether the value of the music spectrum EN counter exceeds 80% of the number of segments. If this check is satisfied, the interval is classified as music at block 102 . Only when all three music tests fail, the logic ends at state 108 without classifying the segment as "music."</p><p>Referring now to FIG. 9 , shown is a check herein for determining whether a segment is a lyric (SOM). A DO loop is provided for each segment within the interval. Proceeding to decision block 112, it is determined whether the percentage of frequencies in the segment within the third frequency band R3 is greater than 15% relative to all sampled frequencies within the segment under examination. If so, the SOM frequency counter is uniformly incremented at block 114 .</p><p>From block 114 or decision block 112, if the test is negative, logic moves to decision block 116 to determine whether the zero crossing rate (ZCR) of the segment under test is greater than or equal to 5 or less than 10. . If so, the SOM ZCR counter is uniformly incremented at block 118 . Logic from block 118 or decision block 116 to decision block 120 if the test is negative to determine whether the normalized third spectral energy concentration NR3 of the segment under test exceeds 90,000 goes on If so, the SOM spectrum EN counter is uniformly incremented at block 122 . From block 122 or decision block 12, if the check is negative, logic proceeds to decision block 124 to determine whether the last segment in the interval has been checked, otherwise, to get the next segment Therefore, logic moves to block 126 to return to decision block 112 . When the latest segment is examined, the logic proceeds to FIG. 10 .</p><p>After checking the segment spacing, in order to classify the interval as "lyric", one of two situations must be met, i.e. one of their combinations. Beginning at decision block 128, it is determined whether the SOM ZCR counter exceeds 70% of the number of segments in the interval. If so, logic moves to block 130 to classify the interval as "lyric" and index and end the interval. However, if the segment fails the first check at decision block 128, then logic proceeds to decision block 132 to apply a first subtest of the second combination check for lyrics. Specifically, at decision block 132, the logic determines whether the value of the SOM frequency counter is less than 50% of the number of segments in the interval under examination. If the first subtest is satisfied, the logic proceeds to a second subtest at decision block 134 to determine whether the value of the SOM ZCR counter exceeds 15% of the number of segments in the interval. If this subtest is positive, the logic moves to decision block 136 to determine whether the value of the SOM spectral EN counter does not exceed 10% of the number of segments. Only when all three subtests according to the second combinatorial test are satisfied, the logic moves to block 130 to classify the interval as lyric, and one of the subtests in decision blocks 132 , 134 , 136 is If the subtest fails, the logic ends at state 138 without classifying the interval as housekeeping. Any interval that is not classified as quiet, voice, music or lyrics is classified as pending before remembering the interval.</p><p>As described above with respect to FIG. 2 , when the intervals of the audio stream are classified, the intervals of the temporary ordered group are stored in a meta-prestored manner to determine whether the group matches one of the meta-patterns. Matches against the pattern type. The audio stream is then indexed again based on the meta-pattern. 11 illustrates how a user can navigate an audio stream as it is indexed to summarize streams, skim streams, and recapitulate streams.</p><p>Beginning at block 140, a user requirement is received in an audio stream. At block 142, the requested portion of the audio stream is retrieved in response to the user requirement and using the index of the audio stream generated as described above. For example, the user wants to access educational audio that is not commercially available, and only the portion of the audio stream that satisfies the educational meta-pattern is returned to block 144 . In other words, the interval or intervals that satisfy the requirement and/or index thereof are returned at block 144 in any order.</p><p>It should be understood that an index of the audio stream may be provided to block 144, for example, in response to a user requirement to summarize the audio stream. The presentation of this list is a summary of the audio stream. Using the index, users can scroll through gaps in the audio stream, and they can listen to the stream and select what they want to scheme and/or summarize.</p><p>In addition to the methods described above for indexing audio streams, FIGS. 12 and 13 show that additional methods can be used to index audio, in particular by qualitatively defined events of interest within audio events that have been classified as speech. it will be shown Beginning at block 146 of FIG. 12, a change in pitch in the audio stream having voice therein is detected. Following a first logical branch, the method proceeds to block 148 for inputting speech into a speech recognition system such as that described in U.S. Patent No. 5,293,584, assigned to the assignee of the present invention, which is incorporated herein by reference. Move. Proceeding to block 150, the output-word-of the speech recognition system is used to index the audio stream.</p><p>Instead of indexing the audio stream into word content at block 150, the logic from block 146 connects a second branch to block 152, where the event of interest in the speech is described below with respect to FIG. identified. What to do with the event of interest in the voice and how to continue checking for the event of interest is determined qualitatively. 12 , events of interest may include emphasis in speech, hesitation in speech, and conclusion ideas in speech.</p><p>Moving to block 154, the audio stream when the stream contains speech is indexed into a meta-pattern set by the sequence of event intervals of interest. An example of such a meta-pattern is the event-of-interest meta-pattern described below of important ideas set by a sequence of 3 seconds (or more) of strong intervals. And, at block 156, the audio stream may also be indexed based on individual events of interest therein.</p><p>Referring now to FIG. 13 , a method for determining the presence of three preferred events/metapatterns of interest is illustrated. Beginning at block 160, a sample of the audio stream is obtained. In one preferred embodiment, each sample has a duration of 10 milliseconds.</p><p>Proceeding to block 162, a dominant frequency of each sample is determined. In determining the dominant frequency, the presently preferred embodiment considers the following eight frequency bands.</p><p>R1-100Hz to 3,000Hz, R2-3,000Hz to 4,000Hz,</p><p>R3-4,000Hz to 5,000Hz, R4-5,000Hz to 6,000Hz,</p><p>R5-6,000Hz to 6,500Hz, R6-6,500Hz to 7,000Hz,</p><p>R7-7,000 Hz to 7,500 Hz, R8-7,500 Hz to 8,000 Hz.</p><p>For each sample, the dominant frequency is calculated as</p><p>RnFreq = the number of frequencies in the n (n = 1 to 8)th band divided by the total number of samples where the dominant frequency range is defined as the longest one of the (8) values for RnFreq</p><p>Moving to block 164, the dominant frequencies are normalized to a histogram. Once the dominant frequency of the audio stream sample has been determined and normalized, processing moves to block 166 to identify the sample having a dominant frequency within the upper 1% of the frequency limit.</p><p>Branching first to decision block 168, the logic determines whether a given sequence in the audio stream contains more than 100 consecutive samples with predominant frequencies within the lower limit of 1%. It should be understood that shorter or longer periods may be used. If such an order is found, the logic proceeds to block 170 to sort and index the order as the final idea in the speech before terminating at state 172 . Otherwise, at decision block 168 the logical branch ends at state 172 .</p><p>Incidentally, the logic branches to decision block 174, where the logic determines whether a given sequence in the audio stream contains more than 100 consecutive samples with dominant frequencies within an upper limit of 1%. It should be understood that shorter or longer periods may be used. If such an order is found, the logic proceeds to block 176 to classify and index the order as an emphasis in speech before terminating at state 172 . Otherwise, the logical branch in decision block 174 ends at state 172 .</p><p>As shown in Figure 13, if an emphasis order in the speech is found, logic proceeds from block 176 to decision block 178, where it is determined whether the strong order is of a duration of at least 3 seconds. However, shorter or longer durations may be used. If such an extended strong order is found, the logic classifies and indexes the order as an important idea language at block 180 . From block 180 and decision block 178 , when the test is negative, the logic ends at state 172 .</p><p>Moreover, the qualitative guidance herein for determining events of interest in speech may include taking into account the rate of change of pitch, amplitude and rate of change of amplitude, as well as other sound characteristics.</p><p>FIG. 14 shows the logic for determining whether an interval is preferably a harmonic frequency for use in the test described above of FIG. 6 . Beginning at block 200, a DO loop is provided for each segment within the interval. Moving to decision block 202, it is determined whether the order of the last frequencies fR is the same as the order of the last frequencies fR for the segment just progressing.</p><p>With respect to the final frequency fR as seen in the present invention, the frequency f1 maintains the following relation, f2 = (I/I+1) * f1 (here, I is an integer greater than or equal to 2) to be true. has at least one musical harmonic frequency f2. When f1 and f2 are provided simultaneously, the final frequency fR is given by the relation fR = f1/I. This is the final frequency fR used in the check in decision block 202 .</p><p>If the check at decision block 202 is negative, the logic moves to decision block 204 to determine whether the last segment has been tested, otherwise the logic searches for the next segment at block 206 . In other words, when the check at decision block 202 is positive, the logic proceeds to block 208 to name the segment under test as a harmonic.</p><p>When the last segment has been examined, logic proceeds from decision block 204 to decision block 210 . At decision block 210, it is determined whether the predetermined sequence of harmonic segments is at least equal to a predetermined period of time, eg, 2 seconds. Otherwise, the logic ends at state 212 . Otherwise, the interval is named as a harmonic in block 214 for use in, for example, the examination of FIG. 6 .</p>
<p>The specific automatic audio content analysis system and method for word spotting, indexing, sorting and retrieving shown and described in detail can fully achieve the above object of the present invention, but since it is a preferred embodiment of the present invention, the present invention The scope of the present invention is limited only by the appended claims, since the scope of the present invention encompasses other embodiments known to those skilled in the art.</p>
17 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| KR20030059503A | Cited by | Republic of Korea | Search report |
| KR100695009B1 | Cited by | Republic of Korea | Search report |
| KR20030070179A | Cited by | Republic of Korea | Search report |
8 members in 5 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 23466399 | United States of America | A | |
| 23466399 | United States of America | A | |
| 9234663 | United States of America | – | |
| 9234663 | – | – | – |
| US19990234663 | – | – | – |
Members8
| Document | Office | Kind | |
|---|---|---|---|
| CN1261181A | China | A | |
| JP2000259168A | Japan | A | |
| KR20000076488AThis record | Republic of Korea | A | |
| US6185527B1 | United States of America | B1 | |
| TW469422B | Taiwan Province of China | B | |
| KR100380947B1 | Republic of Korea | B1 | |
| JP3531729B2 | Japan | B2 | |
| CN1290039C | China | C |
6 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapse due to unpaid annual feeLapsedLAPS | LAPS | |
| Annual fee paymentFPAY | FPAY | |
| Written decision to grantGRNT | GRNT | |
| Decision to grant or registration of patent rightE701 | E701 | |
| Notification of reason for refusalE902 | E902 | |
| Request for examinationA201 | A201 |
Numbers
- Publication
- 1020000076488
- Publication, DOCDB
- 20000076488
- Publication, EPODOC
- KR20000076488
- Application
- 100002364
- Application, DOCDB
- 20000002364
- Application, EPODOC
- KR20000002364
Titles4
- Korean
- 오디오 신호를 분석하기 위한 컴퓨터 구현 방법 및컴퓨터와 그 컴퓨터 프로그램 제품
- English
- Computer implemented method and computer and computer program product for analyzing audio signal
- Unlabeled
- 오디오 신호를 분석하기 위한 컴퓨터 구현 방법 및 컴퓨터와 그 컴퓨터 프로그램 제품{SYSTEM AND METHOD FOR AUTOMATIC AUDIO CONTENT ANALYSIS FOR WORD SPOTTING, INDEXING, CLASSIFICATION AND RETRIEVAL}
- Unlabeled
- A computer implemented method for analyzing an audio signal, and a computer and a computer program product thereof
Classification
- CPC, 7
- G10L15/26
- G10H2210/046
- G10H2210/061
- G10L2015/088
- G06F16/64
- G06F16/685
- G10L25/48
- IPC, 10
- G06F3 16
- G06F17 30
- G10L15 00
- G10L15 04
- G10L15 08
- G10L15 10
- G10L15 18
- G10L15 26
- G10L17 26
- G10L25 00