Interactive voice recognition system and method
Abstract
The present invention relates to an interactive voice recognition system and method, in which voice data of a scenario net structure is recorded in a memory to detect a scenario net number according to a user's voice command, and the detected scenario net By reading the voice data recorded in the memory by number and converting the read voice data into an audible frequency and outputting it, the voice data desired by the user can be provided to the user by step-by-step inquiry and response between the user and the system. It is effective in enabling an interactive conversation between the user and the product.Speech, Recognition, Interactive

Term
Term ended
Expired 25 March 2023, 3.5 years ago.
- Priority and filed
- Granted
- Expired
- Today
23 claims: 7 independent, 16 dependent
- 1삭제
- 2시나리오 넷 구조로 기록된 음성 데이터를 독출하기 위한 제 1 메모리;음성인식용 연산 정보가 미리 기록되어 있고, 사용자의 기록명령에 의해 적어도 하나 이상의 메시지 정보 또는 스케줄 정보가 기록되며, 사용자의 재생명령에 의해 해당 메시지 정보 또는 스케줄 정보를 독출하기 위한 제 2 메모리;마이크를 통해 입력된 사용자의 음성 신호를 디지탈 변환하여 출력하는 음성 입력수단;사용자의 조작에 따라 스위칭 신호를 출력하는 수동 조작수단;상기 음성 입력수단으로부터 입력되는 음성 신호 또는 상기 수동 조작수단에 의해 입력되는 스위칭 신호를 인식하여, 인식 결과 기록명령이면 상기 음성 입력수단으로부터 입력된 음성 신호를 상기 제 2 메모리에 기록하고, 인식 결과 재생명령이면 시나리오 넷 번호를 검출하여 상기 검출된 시나리오 넷 번호에 의해 상기 제 1 메모리에 기록된 음성 데이터를 독출하거나 또는 상기 제 2 메모리에 기록된 음성 데이터를 독출하는 음성 처리수단;및 상기 음성 처리수단으로부터 독출된 음성 데이터를 아날로그 변환하여 스피커를 통해 출력하는 음성 출력수단;을 포함하고, 상기 제 1 메모리는 상기 수동 조작수단를 통해 입력되는 스위칭 신호의 입력 횟수에 대한 정보와 상기 입력 횟수에 대응하는 음성 데이터를 더 포함하고 있어서, 상기 수동 조작수단에 의해 스위칭 신호가 인식되는 경우, 상기 음성 처리수단이 상기 수동 조작수단으로부터 입력된 스위칭 신호의 입력 횟수에 따라 상기 제 1 메모리에 기록된 음성 데이터를 독출함을 특징으로 하는 대화형 음성 인식 시스템.
- 3제 2 항에 있어서, 상기 제 1 메모리에는, 사용자와 시스템간에 이루어지는 단계적인 질의 및 응답에 의해 사용자가 원하는 음성 데이터가 독출되도록, 예상되는 사용자의 질의에 대응되는 적어도 하나 이상의 음성 데이터가 일정한 시나리오 넷 구조로 이미 기록되어 있는 것을 특징으로 하는 대화형 음성 인식 시스템.
- 4제 3 항에 있어서, 상기 제 1 메모리는, 메인 어드레스(Main Address) 영역, 서브 어드레스(Sub-Address) 영역 및 실제 정보(Real Information) 영역으로 구분되며, 상기 메인 어드레스 영역은, 음성인식용 정보를 갖는 영역, 압축신장용 정보를 갖는 영역, 시나리오 넷 번호 정보를 갖는 영역, 시나리오 음성 데이터 정보를 갖는 영역, 인식처리 제어용 정보를 갖는 영역, 기타 음성 데이터 정보를 갖는 영역을 포함하고, 상기 서브 어드레스 영역은, 음성 인식정보 어드레스(N 1 )를 갖는 영역, 압축용 신장용 정보 어드레스(N 2 )를 갖는 영역, 기타 정보(N n )를 갖는 영역을 포함하고, 상기 실제 정보 영역은, 적어도 하나 이상의 음성인식용 정보의 실제 데이터를 갖는 영역, 적어도 하나 이상의 압축 신장용 실제 데이터를 갖는 영역, 적어도 하나 이상의 기타 정보의 실제 데이터를 갖는 영역을 포함하는 것을 특징으로 하는 대화형 음성 인식 시스템.
- 5삭제
- 6제 2 항에 있어서, 상기 음성 입력수단은, 마이크를 통해 입력된 아날로그 음성 신호의 크기를 소정 레벨로 조정하여 출력하는 레벨 조정부;및 상기 레벨 조정부로부터 입력된 아날로그 음성 신호를 입력받아 디지털 음성 신호로 변환하여 출력하는 A/D 컨버터를 포함하는 것을 특징으로 하는 대화형 음성 인식 시스템.
- 7제 2 항에 있어서, 상기 음성 처리수단은, 음성인식용 연산 정보에 의해 상기 음성 입력수단으로부터 입력된 음성 신호를 인식하여 인식 정보를 출력하는 음성 인식부;상기 제 1 메모리 또는 상기 제 2 메모리로부터 독출된 음성 데이터를 소정의 데이터 신장 방식에 따라 신장하는 음성 신장부;상기 음성 입력수단으로부터 입력된 음성 신호를 소정 데이터 압축 방식에 따라 압축하여 상기 제 2 메모리에 기록하는 음성 압축부;노말 상태에서 상기 음성 입력수단으로부터 출력된 음성 신호를 상기 음성 인식부로 입력시키다가 스위칭 제어신호가 입력되면 상기 음성 입력수단으로부터 출력된 음성 신호를 상기 음성 압축부에 입력시키는 스위칭부;및 상기 음성 인식부로부터 입력된 인식 정보에 의해 기록명령인가 재생명령인가를 판단하여, 기록명령이면 상기 스위칭부에 스위칭 제어신호를 출력하여 상기 음성 입력수단으로부터 출력된 음성 데이터를 상기 음성 압축부로 입력시키도록 제어하고, 재생명령이면 시나리오 넷 번호를 검출하여 상기 검출된 시나리오 넷 번호에 의해 상기 제 1 메모리에 기록된 음성 데이터를 상기 음성 신장부로 출력하도록 제어하거나 또는 상기 제 2 메모리에 기록된 음성 데이터를 상기 음성 신장부로 출력하도록 제어하는 제어부를 포함하는 것을 특징으로 하는 대화형 음성 인식 시스템.
- 8제 7 항에 있어서, 상기 음성 인식부는, 사용자의 음성 신호 패턴이 달라진다거나 또는 연속어 문장 형태의 음성 신호가 입력되더라도 핵심적인 단어만을 인식할 수 있도록 하기 위해, HMM(Hidden Markove Model)을 이용한 비터비 알고리즘을 사용하여 음성 신호를 음소 단위로 인식하는 것을 특징으로 하는 대화형 음성 인식 시스템.
- 9제 2 항에 있어서, 상기 음성 출력수단은, 상기 음성 처리수단으로부터 입력된 디지털 음성 데이터를 아날로그 음성 신호로 변환하여 출력하는 D/A 컨버터;및 상기 D/A 컨버터로부터 입력된 아날로그 음성 신호를 전력 증폭하여 스피커를 통해 출력하는 전력 증폭부를 포함하는 것을 특징으로 하는 대화형 음성 인식 시스템.
- 10삭제
- 11제 2 항에 있어서, 상기 음성 처리수단의 제어에 따라 상기 제 1 메모리 또는 상기 제 2 메모리로부터 독출된 음성 데이터를 문자 메시지 또는 영상 이미지로 변환하여 화면에 표시하기 위한 디스플레이 구동수단을 더 포함하는 것을 특징으로 하는 대화형 음성 인식 시스템.
- 12제 2 항에 있어서, 상기 음성 처리수단의 제어에 따라 상기 제 1 메모리 또는 상기 제 2 메모리로부터 독출된 음성 데이터에 의해 해당 메커니즘을 구동시키는 메커니즘 구동수단을 더 포함하는 것을 특징으로 하는 대화형 음성 인식 시스템.
- 13사용자에 의해 입력된 음성 신호나 스위치 신호를 인식하여, 음성 신호를 기록할 것인가 아니면 음성 데이터를 재생할 것인가를 결정하는 제 10 단계;상기 제 10 단계에서 음성 신호 기록이 결정되었으면, 음성인식용 연산 정보가 미리 기록되어 있는 제2 메모리에 외부로부터 입력되는 사용자 음성신호인 적어도 하나 이상의 메시지 정보 또는 스케줄 정보를 기록하는 제 20 단계;및 상기 제 10 단계에서 사용자 음성 신호에 의해 음성 신호 재생이 결정되면, 시나리오 넷 번호를 검출하여 상기 검출된 시나리오 넷 번호에 의해 제 1 메모리에 시나리오 넷 구조로 기록된 음성 데이터를 독출하여 가청주파수로 변환 출력하거나 또는 제 2 메모리에 기록된 음성 데이터를 독출하여 가청주파수로 변환 출력하는 제 30 단계;상기 제 10 단계에서 사용자의 스위치 신호에 의해 음성 신호 재생이 결정되면, 사용자에 의해 입력되는 스위치 신호의 입력회수를 체크하여 상기 제 1 메모리에 기록된 스위칭 신호의 입력회수에 대한 정보를 검출하고, 상기 입력회수 정보에 대응하여 상기 제 1 메모리에 기록된 음성 데이터를 독출하여 가청주파수로 변환 출력하는 제 40 단계를 포함하는 것을 특징으로 하는 대화형 음성 인식 방법.
- 14제 13 항에 있어서, 상기 제 1 메모리는, 예상되는 사용자의 질의에 대응되는 적어도 하나 이상의 음성 데이터가 일정한 시나리오 넷 구조로 이미 기록되어 있어, 사용자와 시스템간에 이루어지는 단계적인 질의 및 응답에 의해 사용자가 원하는 음성 데이터가 독출되는 것을 특징으로 하는 대화형 음성 인식 방법.
- 15제 13 항 또는 제 14 항에 있어서, 상기 제 1 메모리는, 메인 어드레스(Main Address) 영역, 서브 어드레스(Sub-Address) 영역 및 실제 정보(Real Information) 영역으로 구분되며, 상기 메인 어드레스 영역은, 음성인식용 정보를 갖는 영역, 압축신장용 정보를 갖는 영역, 시나리오 넷 번호 정보를 갖는 영역, 시나리오 음성 데이터 정보를 갖는 영역, 인식처리 제어용 정보를 갖는 영역, 기타 음성 데이터 정보를 갖는 영역을 포함하고, 상기 서브 어드레스 영역은, 음성 인식정보 어드레스(N 1 )를 갖는 영역, 압축용 신장용 정보 어드레스(N 2 )를 갖는 영역, 기타 정보(N n )를 갖는 영역을 포함하고, 상기 실제 정보 영역은, 적어도 하나 이상의 음성인식용 정보의 실제 데이터를 갖는 영역, 적어도 하나 이상의 압축 신장용 실제 데이터를 갖는 영역 및 적어도 하나 이상의 기타 정보의 실제 데이터를 갖는 영역을 포함하는 것을 특징으로 하는 대화형 음성 인식 방법.
- 16삭제
- 17제 13 항에 있어서, 상기 제 10 단계는, 시스템 전원이 온되면 초기 시나리오 넷 번호를 검색하는 제 11 단계;외부로부터 음성 신호를 입력받는 제 12 단계;상기 제 11 단계에서 검색된 시나리오 넷 번호에 의해 시나리오 넷 번호 정보를 검색하여 상기 검색된 시나리오 넷 번호 정보에 의해 외부로부터 입력된 음성 신호를 인식할 것인가 아니면 기록할 것인가 아니면 상기 제 1 메모리 또는 상기 제 2 메모리에 기록된 음성 데이터를 재생할 것인가를 결정하는 제 13 단계;상기 제 13 단계에서 음성 신호 인식을 결정하였으면, 상기 시나리오 넷 번호 정보에 의해 인식 도메인을 검색하는 제 14 단계;상기 제 14 단계에서 검색된 인식 도메인에 의해 외부로부터 입력된 음성 신호를 인식하는 제 15 단계;상기 제 15 단계에서 인식된 결과를 시나리오 넷 번호 영역에 존재하는 인식용 패턴 정보로 처리하는 제 16 단계;및 상기 제 16 단계에서 처리된 인식 결과에 의해 다음 시나리오 넷 번호를 검색한 후 상기 제 13 단계로 복귀하는 제 17 단계로 이루어짐을 특징으로 하는 대화형 음성 인식 방법.
- 18제 17 항에 있어서, 상기 제 20 단계는, 상기 제 13 단계에서 음성 신호 기록이 결정되었으면 외부로부터 입력된 음성 신호를 소정의 데이터 압축 방식에 따라 압축시키는 제 21 단계;상기 제 21 단계에서 압축된 음성 데이터를 제 2 메모리에 기록하는 제 22 단계;및 상기 제 16 단계에서 처리된 인식 결과에 의해 다음 시나리오 넷 번호를 검색한 후 상기 제 13 단계로 복귀하는 제 23 단계로 이루어짐을 특징으로 하는 대화형 음성 인식 방법.
- 19제 17 항에 있어서, 상기 제 30 단계는, 상기 제 13 단계에서 음성 데이터 재생이 결정되었으면, 상기 제 13 단계에서 검색된 시나리오 넷 번호 정보에 의해 상기 제 1 메모리에 기록된 해당 음성 데이터를 독출하거나 또는 상기 제 2 메모리에 기록된 음성 데이터를 독출하는 제 31 단계;상기 제 31 단계에서 독출된 음성 데이터를 소정의 데이터 신장 방식에 따라 신장시켜 가청 주파수로 변환 출력하는 제 32 단계;상기 제 16 단계에서 처리된 인식 결과에 의해 다음 시나리오 넷 번호가 존재하는가를 판단하는 제 33 단계;및 상기 제 33 단계에서 판단된 결과 다음 시나리오 넷 번호가 존재하면 다음 시나리오 넷 번호를 검색하여 상기 제 12 단계로 복귀하고, 다음 시나리오 넷 번호가 존재하지 않으면 종료하는 제 34 단계로 이루어짐을 특징으로 하는 대화형 음성 인식 방법.
- 20제 17 항에 있어서, 상기 제 15 단계는, 사용자의 음성 신호의 패턴이 달라진다거나 또는 연속어 문장 형태의 음성 신호가 입력되더라도 핵심적인 단어만을 인식할 수 있도록 하기 위해, HMM(Hidden Markove Model)을 이용한 비터비 알고리즘을 사용하여 음성 신호를 음소 단위로 인식하는 것을 특징으로 하는 대화형 음성 인식 방법.
- 21삭제
- 22제 13 항에 있어서, 상기 제 1 메모리 또는 상기 제 2 메모리로부터 독출된 음성 데이터를 문자 메시지 또는 영상 이미지로 변환하여 화면에 표시하는 제 50 단계를 더 포함하는 것을 특징으로 대화형 음성 인식 방법.
- 23제 13 항에 있어서, 상기 제 1 메모리 또는 상기 제 2 메모리로부터 독출된 음성 데이터에 의해 해당 메커니즘을 구동시키는 제 60 단계를 더 포함하는 것을 특징으로 하는 대화형 음성 인식 방법.
Independent claims23
21 paragraphs in 1 section, as filed
INTERACTIVE VOICE RECOGNITION SYSTEM AND METHOD
1 is a block diagram illustrating an embodiment of an interactive voice recognition system according to the present invention;
2 is a structural diagram of a first memory according to the present invention;
3 is an operation flowchart for explaining a method of reading voice data from a first memory according to the present invention;
4 is a table showing the contents of one scenario net number information according to the present invention;
5 is a configuration diagram of a scenario net recorded in a first memory when the interactive voice recognition system according to the present invention is applied to a watch;
6 is a flowchart illustrating an embodiment of an interactive voice recognition method according to the present invention;
7 is a structural diagram of a first memory having information of a manual operation means according to the present invention;
8 is an operation flowchart for searching for a scenario net according to the operation of the manual operation means according to the present invention.
*Explanation of symbols for main parts of the drawing*
10 : first memory (ROM) 20: second memory (RAM)
30 : voice input means 40: voice processing means
50 : audio output means 60 : manual operation means
70 : display driving means 80: mechanism driving means
31 : Level adjustment unit 32 : A/D converter
41 : voice recognition unit 42: voice extension unit
43 : audio compression unit 44: switching unit
45 : Control unit 51 : D/A converter
52 : Power amplification unit
<backgroundart><p>The present invention relates to an interactive voice recognition system and method, and more particularly, it is possible to provide a user with the voice data desired by the user through a step-by-step query and response between the user and the system by configuring a specific scenario for each product in a memory. It relates to an interactive voice recognition system and method for enabling an interactive conversation between a user and a product.</p><p>Currently, a lot of research and development is being done on a speech recognition system in the form of computer-based software, and there is a trend of continuous development of technology from the existing limited word recognition to continuous word recognition and speech synthesis.</p><p>However, since the conventional computer-based voice recognition system requires expensive computer hardware, there is a problem in that it is very limited in terms of price or size to be applied to overall industrial fields or real life products.</p><p>In addition, the conventional computer-based speech recognition system has made it possible from word-oriented recognition to sentence-based recognition due to active investment in research and development and commercialization. It stays at the level of awareness.</p><p>Looking at the technology in which the conventional voice recognition system is chipped, the user's voice command pattern is recorded within the recognition level range of several tens of words, and whether the user's voice command pattern matches the previously recorded user's voice command pattern when the user makes a voice command It is designed to respond or operate according to the voice command only when it matches after checking.</p><p>Therefore, the conventional voice recognition system has a problem in that it is a speaker-dependent type and does not satisfy the commercial condition that a product to be used in daily life should be convenient, and is sensitive to the user's voice pattern.</p></backgroundart><abstractproblem><p>An object of the present invention to solve the above problems is to record voice data corresponding to an expected user's question in a memory in the form of a predetermined scenario, and then read the voice data from the memory according to the user's voice command. To provide a voice recognition system and a method for controlling the same to provide to the user.</p><p>Accordingly, an object of the present invention is to detect a scenario net number according to a user's voice command, as voice data of a scenario net structure is recorded in a memory, and to retrieve the voice data recorded in the memory according to the detected scenario net number. This is achieved by converting and outputting the read voice data to an audible frequency.</p><p>Another object of the present invention is to provide a 10th step of recognizing a user's voice signal input from the outside and determining whether to record the voice signal or reproduce the voice data; a twentieth step of recording a user voice signal input from the outside in a second memory if it is determined in the tenth step to record the voice signal; If it is determined in the tenth step to reproduce the voice signal, the scenario net number is detected, and the voice data recorded in the scenario net structure in the first memory is read out according to the detected scenario net number, and converted into an audible frequency and outputted in the second memory. This is achieved by performing the thirtieth step of reading out the audio data recorded in the memory and converting it into an audible frequency.</p></abstractproblem>
<p>Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings.</p><p>1 is a block diagram illustrating an embodiment of an interactive voice recognition system according to the present invention.</p><p>As shown in FIG. 1, the present invention includes: a first memory 10 for reading voice data recorded in a scenario net structure; a second memory 20 for recording or reading voice data input by the user's intention; a voice input means 30 for digitally converting and outputting a user's voice signal input through a microphone; Recognizes the voice signal input from the voice input means 30, and records the voice signal input from the voice input means 30 in the second memory 20 if it is a recognition result recording command, and if it is a recognition result reproduction command Voice processing means for detecting the scenario net number and reading the audio data recorded in the first memory 10 or the audio data recorded in the second memory 20 according to the detected scenario net number ( 40); and an audio output means 50 for analog-converting audio data read from the audio processing means 40 and outputting it through a speaker.</p><p>Here, in the first memory 10, at least one or more voice data corresponding to an expected user's query has already been recorded in a predetermined scenario net structure, so that the user's desired voice is generated by a step-by-step query and response between the user and the system. data is read.</p><p>In addition, in the second memory 20, at least one message information or schedule information is recorded by a user's record command, and the corresponding message information or schedule information is read by the user's playback command.</p><p>In the second memory 20, operation information for voice recognition for recognizing a user's voice signal has already been recorded, and an arithmetic space necessary for voice recognition and processing is provided. </p><p>Of course, driving power must be supplied to the audio data recording area of the second memory 20 so that the audio data is not deleted even when the system is down.</p><p>The second memory 20 includes a circular buffer, and temporarily records the digitally converted audio sampling signal in frame units.</p><p>The first memory 10 may be implemented as a ROM, and the second memory 20 may be implemented as a RAM. </p><p>In addition, the audio input means 30 includes: a level adjusting unit 31 for adjusting the size of the analog audio signal input through the microphone to a predetermined level and outputting it; and an A/D converter 32 that receives the analog audio signal input from the level adjusting unit 31, converts it into a digital audio signal, and outputs it.</p><p>In addition, the voice processing unit 40 includes: a voice recognition unit 41 for recognizing the voice signal input from the voice input unit 30 by the operation information for voice recognition and outputting recognition information; an audio expansion unit (42) that expands the audio data read from the first memory (10) or the second memory (20) according to a predetermined data expansion method; an audio compression unit (43) for compressing the audio signal input from the audio input means (30) according to a predetermined data compression method and recording it in the second memory (20); In a normal state, when the voice signal output from the voice input unit 30 is input to the voice recognition unit 41 and a switching control signal is input, the voice signal output from the voice input unit 30 is a switching unit 44 for inputting the audio compression unit 43; It is determined whether a recording command or a reproduction command is received based on the recognition information input from the voice recognition unit 41 , and if it is a recording command, a switching control signal is outputted to the switching unit 44 and outputted from the voice input unit 30 . control to input the voice data into the voice compression unit 43, and if it is a playback command, detect a scenario net number, and convert the voice data recorded in the first memory 10 according to the detected scenario net number to the voice and a control unit 45 for controlling output to the decompression unit 42 or outputting audio data recorded in the second memory 20 to the audio decompression unit 42 .</p><p>Here, the voice recognition unit 41 recognizes a voice signal in phoneme units using a Viterbi algorithm using a Hidden Markove Model (HMM), so that the user's voice signal pattern is changed or a voice signal in the form of a continuous sentence is input. However, only key words can be recognized.</p><p>In addition, the audio output means 50 includes: a D/A converter 51 for converting digital audio data input from the audio processing means 40 into analog audio signals and outputting them; and a power amplifying unit 52 for power amplifying the analog audio signal input from the D/A converter 51 and outputting it through a speaker.</p><p>In addition, the present invention further includes a manual operation means (60) for outputting a switching signal according to a user's operation, and the first memory (10) is the number of times of input of the switching signal input through the manual operation means (60) and voice data corresponding to the number of inputs, so that the voice processing unit 40 controls the first memory 10 according to the number of times the switching signal input from the manual manipulation unit 60 is input. Read the audio data recorded in the .</p><p>In addition, the present invention is a display for converting the audio data read from the first memory 10 or the second memory 20 into a text message or a video image under the control of the audio processing means 40 to display on the screen It further includes a driving means (70).</p><p>In addition, the apparatus further includes a mechanism driving means 80 for driving a corresponding mechanism by the audio data read from the first memory 10 or the second memory 20 under the control of the audio processing means 40 .</p><p>2 is a structural diagram of a first memory according to the present invention.</p><p>As shown in FIG. 2 , the structure of the first memory 10 includes a main address area 11 , a sub-address area 12 , and a real information area ( 13).</p><p>The main address area 11 and the sub address area 12 may be integrated.</p><p>Here, the main address area 11 is subdivided into areas such as speech recognition information, compression/decompression information, scenario net number information, scenario voice data information, recognition processing control information, other voice data information, and other information. </p><p>The voice recognition information designates a first address of a sub-address representing information on parameter values required for voice recognition.</p><p>The compression decompression information designates a first address of a sub-address representing information on parameter values required for compression decompression.</p><p>The scenario net number information is actually created by configuring a scenario net corresponding to each product, and the scenario net is configured so that the user and the voice recognition system can interactively communicate in an area related to the product.</p><p>Each scenario net number information includes all information such as whether to self-recognize, whether to make a fixed response, or whether to move on to the next scenario after responding.</p><p>The scenario voice data information means actual voice data for the scenario net number.</p><p>The other voice data information is voice data information used when a fixed response or values processed by the control unit 45 need to be responded.</p><p>The other information means scenario information added by the manual operation means 60 .</p><p>As such, the main address 11 designates the first address of each sub-address stage, and the sub-address 12 designates different numbers of the first addresses of the actual address stage.</p><p>That is, if the number of scenario net number information is N, the number of scenario net number information in the sub-address stage is also N. And the first address of the actual information of the scenario net number a is in the scenario net number information of the sub-address a-th.</p><p>By configuring the first memory 10 in this way, the efficiency of memory usage is increased.</p><p>The real information area 13 is the area designated by the sub-address 12, such as at least one or more real data of speech recognition information, at least one or more real data of compression/decompression information, and real data of at least one or more other information. have real data.</p><p>Since the compressed audio data has a variable length with different lengths, length information is recorded at the first address of the actual data and the actual audio data is recorded thereafter, thereby increasing the efficiency of memory usage.</p><p>At this time, the actual data is recorded in the form of a wave file.</p><p>3 is an operation flowchart illustrating a method of reading voice data from a first memory according to the present invention.</p><p>As shown in FIG. 3 , when a voice command is input from the user or when a switching signal corresponding to the voice command is input through the manual operation means 60 , the controller 45 first searches the main address area. (S51), a sub-address area is searched using the search result of the main address area (S52), a real address area is searched using the search result of the sub-address area (S53), and the real address area is searched for (S53). The actual data is extracted using the search result (S54).</p><p>That is, in the main address, the first address of the sub-address end is searched according to the type of information, and the first address of the data to be actually known is searched according to the number of the sub-address end.</p><p>4 is a table showing the contents of one scenario net number information according to the present invention.</p><p>One scenario net number information includes information for fixed data, information for recognition, domain number information, next scenario net number information, information on the number of patterns for recognition, and scenario net number information, as shown in FIG. 4 . </p><p>The fixed data information includes fixed response information of the scenario net number, that is, information that responds with a result obtained by an algorithm in the control unit 45 .</p><p>The recognition information includes the recognized processing result information, the domain number information includes domain number information of a target to be recognized by the voice recognition unit 41, and the pattern number information includes pattern number information of each recognition target. has a</p><p>In addition, information for a pure scenario response may be included in the domain number information.</p><p>The scenario net number information has a plurality of scenario net numbers in which the next scenario is developed, so that random processing is possible, thereby avoiding a simple configuration. </p><p>Finally, since the pattern information for recognition and the scenario net number information have as many as the number of patterns, it is necessary to determine the recognized words and sentences by providing dictionary information consisting of words or sentences according to the pattern or voice recognition method. have information about</p><p>All such scenario information and system control processing information are configured in the first memory 10 .</p><p>5 is a block diagram of a scenario net recorded in the first memory 10 when the interactive voice recognition system according to the present invention is applied to a watch.</p><p>As shown in FIG. 5 , the interactive voice recognition system according to the present invention maintains a standby state until a voice command for N functions is input after giving an introduction response when the system power is turned on.</p><p>At this time, the response to the user command of each function is already configured as a scenario, each has a scenario net number, and all of this individual information is recorded in the first memory 10 .</p><p>As shown in FIG. 5 , the voice recognition system according to the present invention does not end with only one response to one user's voice command, but the plurality of response scenarios are organically coupled to each other, so that the user and the system By repeating questions and answers in a conversational format, it is possible to provide voice information desired by the user.</p><p>For example, when the user inputs a voice command of "record message", the system according to the present invention outputs a voice response "what message will you record", and accordingly the user inputs the voice command "wake up time" again Then, the system outputs a voice response saying "Start recording", and accordingly, when the user inputs a voice command of "00:00" again, the system records it in the second memory 20 and then says "Recording is complete" ", output the voice response.</p><p>In addition, when the user inputs a voice command "read message", the system according to the present invention outputs a voice response "what message will you read", and accordingly the user inputs a voice command "wake up time" again Then, the system reads the corresponding voice data from the second memory 20 and outputs a voice response of "00: 00".</p><p>In addition, when the user inputs a voice command "time", the system according to the present invention outputs a voice response "what city do you want to know the time of", and accordingly, when the user inputs a voice command "New York" again, The system outputs a voice response saying "The current New York time is "00:00".</p><p>In addition, it has a time-out scenario in which the system is forcibly processed when it is processed as a failure and when no voice is heard for a predetermined period of time, so that the user's voice command can be smoothly processed without interruption. can</p><p>Next, the operation of the interactive voice recognition system according to the present invention configured as described above will be described with reference to FIG. 1 .</p><p>First, when the system power is turned on by the user, the control unit 45 reads the voice data for the method of using the voice recognition system recorded in the first memory 10 and outputs it to the voice expansion unit 42 .</p><p>The audio expansion unit 42 expands the audio data input from the first memory 10 according to a predetermined data expansion method and outputs the expanded audio data to the D/A converter 51 .</p><p>Accordingly, the D/A converter 51 converts the digital audio data input from the audio expansion unit 42 into an analog audio signal and outputs it to the power amplifier 52, and the power amplifier 52 By amplifying the power of the analog audio signal input from the D/A converter 51 and outputting it through a speaker, the user is informed of basic usage for a predetermined time and ready to start.</p><p>As described above, the user who is familiar with the method inputs a voice command to obtain desired information. </p><p>Accordingly, when the user's voice command is input to the voice input means 30 through the microphone, first, the level adjusting unit 31 adjusts the size of the analog voice signal input through the microphone to a predetermined level to the A/D converter 32 . output as</p><p>The A/D converter 32 receives the analog audio signal input from the level adjusting unit 31 , converts it into a digital audio signal, and outputs it to the audio processing unit 40 .</p><p>The voice processing means 40 recognizes the voice signal input from the A/D converter 32, and if it is a recognition result recording command, the voice signal input from the A/D converter 32 is converted into the second memory 20 ), and if the recognition result is a reproduction command, the scenario net number is detected and the voice data recorded in the first memory 10 is read or recorded in the second memory 20 according to the detected scenario net number. Read audio data.</p><p>The operation of the voice processing unit 40 will be described in more detail as follows.</p><p>First, the voice recognition unit 41 recognizes the voice signal input from the A/D converter 32 using the operation information for voice recognition recorded in the second memory 20 and outputs recognition information.</p><p>Here, the voice recognition unit 41 recognizes a voice signal in phoneme units using a Viterbi algorithm using a Hidden Markove Model (HMM), so that the user's voice signal pattern is changed or a voice signal in the form of a continuous sentence is input. However, only key words can be recognized.</p><p>For example, when the voice recognition system according to the present invention is applied to a watch, even if the user inputs a voice command in the form of different continuous sentences such as "time" or "what time" or "what time is now" or "what time is now" , The speech recognition system according to the present invention recognizes only the key word "time" by the Viterbi algorithm using HMM (Hidden Markove Model) and provides a voice signal "The current time is 00:00" to the user.</p><p>In order to recognize a voice using the HMM (Hidden Markove Model), the control unit 45 performs numerous calculations. At this time, the necessary constants are recorded in the second memory 20, so that the control unit 45 whenever necessary. is read under the control of</p><p>The control unit 45 uses the second memory 20 to calculate, write, and read a necessary value. Since the calculation of the data is huge, a separate memory management unit is provided for voice recognition. Dedicated to data management functions.</p><p>Accordingly, the control unit 45 determines whether the recognition information input from the voice recognition unit 41 is a recording command or a reproduction command, and outputs a switching control signal to the switching unit 44 if the determination result is a recording command. The audio data output from the A/D converter 32 is controlled to be input to the audio compression unit 43, and if the determination result is a reproduction command, the scenario net number is detected, and the first The audio data recorded in the memory 10 is controlled to be output to the audio expansion unit 42 , or the audio data recorded in the second memory 20 is controlled to be output to the audio expansion unit 42 .</p><p>Accordingly, if the voice command is a recording command, the voice compression unit 43 receives the voice data output from the A/D converter 32 through the switching unit 44 under the control of the control unit 45 and receives a predetermined value. The audio is compressed according to the audio compression method.</p><p>The control unit 45 writes the compressed audio signal to a predetermined address of the second memory 10 . </p><p>On the other hand, if the determination result is a replay command, the control unit 45 detects the scenario net number recorded in the first memory 10 and transfers actual data to the first memory 10 or the first memory 10 using the detected scenario net number. When the data is read from the second memory 20 and the read data is composed of a plurality of different voice signals rather than actual data, the controller 45 may control the user's voice in response to any one of the plurality of voice signals. It is determined whether a command is input, the corresponding actual data is read from the first memory 10 or the second memory 20 , and output to the voice decompression unit 42 . </p><p>Accordingly, under the control of the controller 45, the voice expansion unit 42 receives the voice data recorded in the first memory 10 or the second memory 20 and expands it according to a predetermined voice expansion method. Then, it is output to the audio output means (50).</p><p>The audio data expanded by the audio expansion unit 42 is converted into an analog audio signal in the D/A converter 51 and output to the power amplifier 52, and the power amplifier 51 is the D/A converter. The analog audio signal input from (51) is amplified by power and outputted through a speaker.</p><p>At this time, the control unit 45 determines whether the next scenario net number exists according to the scenario net number information, ends if the next scenario net number does not exist, and searches for the next scenario net number if the next scenario net number exists Thus, by repeating the above processes, the voice data desired by the user can be provided to the user through step-by-step questions and responses made between the user and the system, thereby enabling an interactive conversation between the user and the product. </p><p>Next, referring to FIG. 6 , the flow of the interactive voice recognition method according to the present invention configured as described above is as follows.</p><p>6 is a flowchart illustrating an embodiment of an interactive voice recognition method according to the present invention.</p><p>First, the controller 45 recognizes the user's voice signal input from the outside in the tenth step (S10), and determines whether to record the voice signal or reproduce the voice data.</p><p>That is, when the system power is turned on by the user, the control unit 45 searches for the net number of the initial scenario in the eleventh step ( S11 ). </p><p>Accordingly, when receiving the user's voice command from the outside in the twelfth step (S12), the control unit 45 searches for the scenario net number information by the searched scenario net number in the thirteenth step (S13), and the searched scenario net The number information determines whether to recognize or record a voice signal input from the outside, or whether to reproduce voice data recorded in the first memory 10 or the second memory 20 .</p><p>When the voice signal recognition is determined in the thirteenth step (S13), the controller 45 searches for a recognition domain number based on the scenario net number information in a fourteenth step (S14), and in a fifteenth step (S15) ), after recognizing a voice signal input from the outside using the searched domain number, in step 16 (S16), the recognition result is processed as pattern information for recognition existing in the scenario net number information area, and then in the 17th step (S16). After searching for the scenario net number of the next process based on the recognition result processed in step S17, the process returns to the thirteenth step (S13).</p><p>On the other hand, when it is determined in the tenth step (S10) to record the voice signal, the control unit 45 records the user voice signal input from the outside in the second memory (20) in the twentieth step (S20).</p><p>That is, when it is determined in the thirteenth step (S13) to record the voice signal, the control unit 45 compresses the user voice signal input from the outside in the twenty-first step (S21) according to a predetermined data compression method, and then performs the 22nd step (S21). After recording the compressed voice data in the second memory 20 in step S22, the next scenario net number is searched for by the scenario net number information in the 23rd step S23, and then in the thirteenth step (S13). return to </p><p>On the other hand, when the scenario reproduction is determined in the tenth step (S10), the controller 45 detects the scenario net number in the thirty (S30) step and stores it in the first memory 10 according to the detected scenario net number. The voice data recorded in the scenario net structure is read and converted into an audible frequency, or the voice data recorded in the second memory 20 is read and converted into an audible frequency and output.</p><p>That is, when it is determined in the thirteenth step (S13) to reproduce the voice data, the controller 45 reads the voice data recorded in the first memory 10 according to the scenario net number information in the 31st step (S31). Alternatively, after reading the voice data recorded in the second memory 20, in the 32nd step (S32), the read voice data is expanded according to a predetermined data expansion method, converted into an audible frequency, and outputted. In step 33 (S33), it is determined whether the next scenario net number exists based on the scenario net number information. It returns to step 12 (S12), and if the next scenario net number does not exist, it ends.</p><p>That is, the control unit 45 determines whether the next scenario net number exists according to the scenario net number information, ends if the next scenario net number does not exist, and searches for the next scenario net number if the next scenario net number exists Thus, by repeating the above processes, the voice data desired by the user can be provided to the user through step-by-step questions and responses made between the user and the system, thereby enabling an interactive conversation between the user and the product. </p><p>For example, the interactive voice recognition method according to the present invention will be described in detail with reference to FIGS. 5 and 6 .</p><p>First, when the system power is turned on by the user, the control unit 45 searches for the net number N0 of the initial scenario in the eleventh step (S11). </p><p>Accordingly, when the user's voice signal "message record" is input in the twelfth step (S12), the control unit 45 uses the initial scenario net number (N0) found in the thirteenth step (S13) to determine the initial scenario net number. The information (N0_data) is retrieved and the recognition of the voice signal is determined based on the retrieved scenario net number information (N0_data). </p><p>The control unit 45 searches for a recognized domain number based on the initial scenario net number information (N0_data) in step 14 (S14), and uses the searched domain number in step 15 (S15) from the outside. After recognizing the input voice signal, in step 16 (S16), the recognition result is processed as pattern information for recognition existing in the scenario net number information area, and then in step 17 (S17), the recognition result is added to the recognition result. After searching for the next scenario net number (N1) by the method, the process returns to the thirteenth step (S13).</p><p>Accordingly, the control unit 45 secures a message recording space in a predetermined location of the second memory 20, and in the thirteenth step (S13), the scenario net number information (N1_data) is stored by the scenario net number (N1). Search and determine the reproduction of voice data based on the searched scenario net number information (N1_data).</p><p>Accordingly, in step 31 (S31), the control unit 45 reads out voice data indicating "message recording start" from a predetermined area of the first memory 10 according to the scenario net number information N1_data, and In step 32 (S32), the read voice data is expanded according to a predetermined data expansion method and converted into an audible frequency, and then the next scenario net number is determined by the scenario net number information (N1_data) in step 33 (S33). It is determined whether (N2) exists, and if the next scenario net number (N2) exists, the next scenario net number (N2) is searched for in step 34 (S34) and the process returns to step 12 (S12).</p><p> Accordingly, the user inputs the message voice to be recorded in accordance with the response voice of "Start message recording".</p><p>Accordingly, when the user's message information is input in the twelfth step (S12), the control unit 45 searches the scenario net number information (N2_data) by the scenario net number (N2) in the thirteenth step (S13), and the Recognition of a voice signal is determined by the searched scenario net number information (N2_data). </p><p>The control unit 45 searches for a recognized domain number based on the scenario net number information (N2_data) in step 14 (S14), and uses the searched domain number in step 15 (S15) to input from the outside. After recognizing the voice signal, the recognition result is processed as pattern information for recognition existing in the scenario net number information area in step 16 (S16), and then, by the recognition result processed in step 17 (S17) After searching for the next scenario net number (N3), the process returns to the thirteenth step (S13).</p><p>Accordingly, the controller 45 retrieves the scenario net number information (N3_data) by the scenario net number (N3) in the thirteenth step (S13) and determines the recording of voice data according to the scenario net number information (N3_data) do.</p><p>Accordingly, the controller 45 compresses the message information input from the outside in the 21st step (S21) according to a predetermined data compression method, and then stores the compressed voice data in the second memory in the 22nd step (S22). After recording in the previously secured area of (20), the next scenario net number (N4) is searched for by the scenario net number information (N3_data) in the 23rd step (S23), and then the thirteenth step (S13) is returned. . </p><p>The controller 45 retrieves the scenario net number information (N4_data) by the scenario net number (N4) in the thirteenth step (S13) and determines the reproduction of the voice data based on the retrieved scenario net number information (N4_data) .</p><p>Accordingly, in step 31 ( S31 ), the control unit 45 reads out voice data indicating "message recording is complete" from the predetermined area of the first memory 10 based on the scenario net number information (N4_data), and In step 32 (S32), the read voice data is expanded according to a predetermined data expansion method, converted to an audible frequency and output, and then in step 33 (S33), the next scenario net number is determined by the scenario net number information (N4_data). It is determined whether (N5) exists, and if the next scenario net number (N5) does not exist, the system is terminated. </p><p>In addition, the voice recognition system according to the present invention includes a manual operation means 60 for outputting a switching signal according to a user's operation as shown in FIG. 1 , and the first memory 10 is the manual operation means It includes information on the number of times of input of the switching signal input through (60) and voice data corresponding to the number of times of input.</p><p>Accordingly, the audio processing means 40 reads the audio data recorded in the first memory 10 according to the number of times the switching signal input from the manual operation means 60 is input.</p><p>The manual operation means 60 is composed of a main switch and a plurality of sub-switches, or is composed of a main switch and a reset switch, and at least one switch is operated by the user to output a switching signal corresponding to the user's voice command. do.</p><p>7 is a structural diagram of a first memory having information of the manual operation means according to the present invention, and FIG. 8 is an operation flowchart for searching for a scenario net according to the operation of the manual operation means according to the present invention.</p><p>As shown in FIG. 7 , the other information area of the first memory 10 may additionally have scenario information added by the manual operation means 60 .</p><p>That is, the other area of the main address 11 has voice data corresponding to the operation of the manual operation means 60 , and the sub address area 12 is a switching signal input from the manual operation means 60 . It has address information of an area in which a corresponding scenario net number is recorded according to the number of times, and the actual information area 13 is subdivided into areas in which at least one or more scenario net numbers corresponding to the manual operation means 60 are recorded. do.</p><p>Therefore, according to the present invention, the number of function addresses is indicated according to the number of times the switch is pressed, so there are as many switches as the number of scenario deployments.</p><p>That is, according to the present invention, the main switch and the sub-switch can be as many as the scenario to be developed by pointing to the function address according to the number of times the main switch is pressed, and when entering a specific function, pressing the sub switch and pointing to the corresponding function address again is used. exist.</p><p>Accordingly, as shown in FIG. 8, the control unit 45 checks how many times the main switch of the manual operation means 60 is pressed by the user (S41), and the main switch is pressed a times by the user. In this case, data written in the a-th region of the sub-address region of the first memory 10 is read, and data written in the real address region of the first memory 10 is read using the data. (S42).</p><p>When the data read from the real address area of the first memory 10 has another response data, the control unit 45 checks how many times the first sub switch is pressed by the user (S43) , when the first sub-switch is pressed b times by the user, data written in the b-th area of the sub-address area of the first memory 10 is read, and the first memory 10 using the data The data recorded in the real address area of ' is read ( S44 ).</p><p>When the data read from the real address area of the first memory 10 has another response data, the control unit 45 checks how many times the second sub-switch is pressed by the user (S45) , when the first sub-switch is pressed c times by the user, data written in the c-th area of the sub-address area of the first memory 10 is read, and the first memory 10 using the data The data recorded in the real address area of ' is read (S46). </p><p>At this time, when the data read from the real address area does not have any more response data, the audio signal is output through the speaker.</p><p>In addition, the voice recognition system according to the present invention includes a display driving means 70 as shown in FIG. 1 , so that the first memory 10 or the second memory according to the control of the voice processing means 40 . The audio data read from (20) is converted into a text message or a video image and displayed on the screen.</p><p>That is, the control unit 45 outputs a display control signal according to the audio data read from the first memory 10 or the second memory 20 in the 50th step (S50), and accordingly the display driving means At 70, a corresponding text message or video image is displayed on the screen by the display control signal. </p><p>In addition, the voice recognition system according to the present invention includes a mechanism driving means 80 as shown in FIG. 1 , so that the first memory 10 or the second memory according to the control of the voice processing means 40 . The mechanism is driven by the voice data read from (20).</p><p>That is, the control unit 45 outputs a mechanism control signal according to the voice data read from the first memory 10 or the second memory 20 in the 60th step (S60), and accordingly, the mechanism driving means ( 80) drives the corresponding mechanism by the mechanism control signal.</p><p>Since all information on the operation flow described above is already recorded in the first memory 10, by replacing only the first memory 10, various functions can be provided to various products.</p>
<p>As described above, the present invention detects the scenario net number according to the user's voice command, since voice data corresponding to the expected user query is already recorded in the memory in a scenario net structure in advance, and the detected scenario net number by reading the voice data recorded in the memory and converting the read voice data into an audible frequency and outputting the voice data desired by the user can be provided to the user through a step-by-step query and response between the user and the system. , it is effective in that not only interactive conversation between the user and the product is possible, but also there is no inconvenience in recording the user first when using the system for the first time.</p><p>In addition, the present invention records message information or schedule information in the second memory according to a user record command, reads the corresponding message information or schedule information from the second memory by a user playback command, and converts the read voice data into an audible frequency. By outputting it, it has the effect of performing the role of a personal assistant.</p><p>In addition, the present invention reads the corresponding voice data recorded in the memory according to the user's switch operation, converts the read voice data into an audible frequency and outputs the output. It works to make it work.</p><p>In addition, the present invention is effective in converting the audio data read from the first memory 10 or the second memory 20 by the display driving means 70 into a text message or a video image and displaying it on the screen. .</p><p>In addition, the present invention is effective in driving the corresponding mechanism by the voice data read from the first memory 10 or the second memory 20 by the mechanism driving means (80).</p><p>In addition, the present invention uses the Viterbi algorithm using HMM (Hidden Markove Model), so that even if the user's voice signal pattern is changed or a voice signal in the form of a continuous sentence is input, only key words can be recognized, so that a natural conversation with the user is possible. The effect is that it is possible.</p><p>In addition, the present invention can be implemented in the form of a chip rather than a computer, and can be applied to various products throughout the industry by combining a voice recognizer in the form of an ASIC chip and a memory recorded in the form of a scenario net. It has an effect.</p><p>In addition, the present invention has an effect that, by replacing only a memory having an algorithm for voice data and voice recognition corresponding to a user's voice command, the system can be easily upgraded and variously applied to various products.</p><p>Examples of application of the present invention include education and entertainment fields (voice recognition interactive toys, e-books, language learners), information and communication fields (voice recognition personal digital assistant, PDA, mobile phone, computer, home automation system, navigator), household goods There are fields (voice recognition watches, stands, televisions, audio, automobiles) and medical fields (voice recognition medical devices for the disabled, silver products).</p><p>Although the above has been described with reference to the preferred embodiment of the present invention, those skilled in the art can variously modify and change the present invention within the scope without departing from the spirit and scope of the present invention as described in the following claims. You will understand that there is</p>
6 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US8888494B2 | Cited by | United States of America | Applicant |
| WO2012006024A3 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US10642463B2 | Cited by | United States of America | Applicant |
| US9122656B2 | Cited by | United States of America | Applicant |
| US9870134B2 | Cited by | United States of America | Applicant |
| WO2012006024A2 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US9904666B2 | Cited by | United States of America | Applicant |
| JP2001249924A | Cites | Japan | Examiner |
| JP2001249924A | Cites | Japan | Search report |
| JPH07239694A | Cites | Japan | Search report |
| JPH09131468A | Cites | Japan | Search report |
| JPH11237971A | Cites | Japan | Search report |
| JPH1185178A | Cites | Japan | Search report |
2 members in 1 office
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 20030018622 | Republic of Korea | A | |
| KR20030018622 | – | – | – |
Members2
| Document | Office | Kind | |
|---|---|---|---|
| KR20040083919A | Republic of Korea | A | |
| KR100554397B1This record | Republic of Korea | B1 |
6 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapse due to unpaid annual feeLapsedLAPS | LAPS | |
| Annual fee paymentFPAY | FPAY | |
| Written decision to grantGRNT | GRNT | |
| Decision to grant or registration of patent rightE701 | E701 | |
| Notification of reason for refusalE902 | E902 | |
| Request for examinationA201 | A201 |
Numbers
- Publication
- 10-0554397
- Publication, DOCDB
- 100554397
- Publication, EPODOC
- KR100554397B
- Application
- 100018622
- Application, DOCDB
- 20030018622
- Application, EPODOC
- KR20030018622
Titles2
- Korean
- 대화형 음성 인식 시스템 및 방법
- English
- Interactive speech recognition system and method
Classification
- CPC, 4
- G10L15/22
- G10L15/08
- G10L2015/225
- G10L2015/221
- IPC, 1
- G10L15 22