Robot apparatus and control method thereof
Abstract
This record has no abstract on file.
Term
Term ended
Expired 22 October 2022, 3.9 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
7 claims: 3 independent, 4 dependent
- 1入力情報に対応して動作するロボット装置であって、 外部環境 から 認識属性毎に特徴量を認識する複数の認識手段と、 前記複数の認識手段からの認識結果を基にターゲットを検出するターゲット検出手段と、 前記ターゲット検出手段により検出されたターゲット 毎 の 複数の 認識属性の 各々についての 特徴量と時間的及び空間的情報を記憶 可能な ターゲット・メモリと、 前記ターゲット・メモリに記憶された各ターゲットに基づいて、前記ロボット装置の行動を決定する行動決定手段と、を具備し、 前記ターゲット検出手段は、 前記複数の認識手段のうちいずれかからの認識結果について、前記ターゲット・メモリ内の時間的及び空間的関係が近いターゲットを同定し、 前記の認識結果が前記の同定したターゲットの同じ認識属性の特徴量と一致するときには、前記の同定した ターゲットの認識属性の情報 に前記 の認識結果 を 反映させる とともに時間的及び空間的情報を更新し 、 前記の同定したターゲットが前記の認識結果と同じ認識属性を持たないときには、前記の同定したターゲットに前記の認識結果を追加するとともに時間的及び空間的情報を更新し、 前記の認識結果が該同定したターゲットの同じ認識属性の特徴量と一致しないとき、又は、前記の認識結果と時間的及び空間的関係が近いターゲットが前記ターゲット・メモリ内にないときには、前記の認識結果を持つ 新規のターゲット を 前記ターゲット・メモリに記憶する、ことを特徴とするロボット装置。
- 2前記複数の認識属性は、少なくとも色情報、声情報、顔認識情報 のうち1つを含む 、ことを特徴とする請求項1に記載のロボット装置。
- 3前記ターゲット検出手段が検出したターゲットの空間的情報を、前記認識手段において使用される座標系から前記ロボット装置を基準にした固定座標系の表現に変換する、ことを特徴とする請求項1に記載のロボット装置。
- 4前記ターゲット・メモリ内に記憶されているターゲットの確信度を時間の経過に従ってデクリメントし、確信度が所定値を下回ったターゲットを前記ターゲット・メモリから削除するガーベッジ・コレクタをさらに備える、ことを特徴とする請求項1に記載のロボット装置。
- 5外部環境で発生したイベントを時系列に応じて、関連するターゲット情報と併せて記憶するイベント・メモリをさらに備え、 前記行動決定手段は、前記ターゲット・メモリに記憶されたターゲットと前記イベント・メモリに記憶されたイベントに基づいて行動決定を行なう、ことを特徴とする請求項1に記載のロボット装置。
- 6前記イベントには、前記ターゲットの出現並びに消失、音声認識単語、対話、前記ロボット装置の自己行動結果のいずれかが含まれている、ことを特徴とする請求項 5 に記載のロボット装置。
- 7入力情報に対応して動作するロボット装置の制御方法であって、 外部環境 から 認識属性毎に特徴量を認識する複数の認識ステップと、 前記複数の認識ステップにおける認識結果を基にターゲットを検出するターゲット検出ステップと、 前記ターゲット検出ステップにより検出されたターゲット 毎 の 複数の 認識属性 各々についての 特徴量と時間的及び空間的情報を記憶するターゲット記憶ステップと、 前記ターゲット検出ステップにより記憶された各ターゲットに基づいて、前記ロボット装置の行動を決定する行動決定ステップと、を有し、 前記ターゲット検出ステップでは、 前記複数の認識ステップのうちいずれかからの各認識属性の認識結果について、前記ターゲット・メモリ内の時間的及び空間的関係が近いターゲットの認識属性を同定し、 前記の認識結果が前記の同定したターゲットの同じ認識属性の特徴量と一致するときには、前記の同定した ターゲットの認識属性の情報 に前記 の認識結果 を 反映させる とともに時間的及び空間的情報を更新し 、 前記の同定したターゲットが前記の認識結果と同じ認識属性を持たないときには、前記の同定したターゲットに前記の認識結果を追加するとともに時間的及び空間的情報を更新し、 前記の認識結果が該同定したターゲットの同じ認識属性の特徴量と一致しないとき、又は、前記の認識結果と時間的及び空間的関係が近いターゲットが前記ターゲット・メモリ内にないときには、前記の認識結果を持つ 新規のターゲット を 前記ターゲット・メモリに記憶する、ことを特徴とするロボット装置の制御方法。
Independent claims7
1 paragraph, as filed
[Technical field] The present invention relates to a robot device that performs autonomous movements and realizes realistic communication and a control method thereof, and in particular, a function that recognizes information in the outside world such as images and sounds and reflects its own actions on it. The present invention relates to an autonomous robot device equipped with the above and a control method thereof. More specifically, the present invention relates to a robot device that integrates a plurality of recognition results from the outside world and treats them as meaningful symbol information and controls their behavior, and particularly to previously observed recognition results. The present invention relates to a robot device and its control method for solving which skin color region corresponds to which person on the face and which person's voice is the voice of which person by using more complicated recognition results such as the correspondence problem of. [Background technology] A mechanical device that uses electrical or magnetic action to perform movements that resemble human movements is called a "robot." The etymology of robots is said to be derived from the Slavic word "ROBOTA (slave machine)". In Japan, robots began to spread from the end of the 1960s, but most of them are industrial robots such as manipulators and transfer robots for the purpose of automating and unmanned production work in factories. Met. Recently, it imitates the body mechanism and movements of quadrupedal animals such as dogs, cats, and bears, and pet-type robots that imitate their movements, or the body mechanisms and movements of humans, monkeys, and other animals that walk upright on two legs. Research and development on the structure of leg-type mobile robots such as "humanoid" or "humanoid robots" and their stable walking control have progressed, and expectations for their practical application are increasing. Compared to crawler robots, these leg-type mobile robots are more unstable and difficult to control posture and walking, but they are superior in that they can realize flexible walking and running movements such as climbing stairs and overcoming obstacles. There is. Stationary type robots, such as arm-type robots, which are planted and used in a specific place, operate only in a fixed or local work space such as parts assembly and sorting work. On the other hand, a mobile robot has a non-limited work space, and can freely move on a predetermined path or a non-route to perform a predetermined or arbitrary human work, or a human or a dog. Alternatively, it can provide various services that replace other life forms. One of the uses of the leg-type mobile robot is to act for various difficult tasks in industrial activities and production activities. For example, maintenance work at nuclear power plants, thermal power plants, petrochemical plants, parts transportation / assembly work at manufacturing plants, cleaning at high-rise buildings, rescue at fire sites, etc. .. In addition, as another use of the leg-type mobile robot, there is a life-oriented type, that is, a use of "coexistence" or "entertainment" with human beings, rather than the above-mentioned work support. This type of robot faithfully reproduces the movement mechanism of relatively intelligent legged walking animals such as humans, dogs (pets), and bears, and rich emotional expressions using limbs. In addition to simply faithfully executing the pre-entered motion pattern, it dynamically responds to words and attitudes (such as "praise", "scolding", and "tapping") received from the user (or other robot). It is also required to realize a corresponding and lively response expression. In the conventional toy machine, the relationship between the user operation and the response operation is fixed, and the toy operation cannot be changed according to the user's preference. As a result, the user eventually gets tired of the toy that repeats only the same operation. On the other hand, an intelligent robot that performs autonomous movement generally has a function of recognizing information in the outside world and reflecting its own behavior on it. That is, the robot realizes autonomous thinking and motion control by changing the emotion model and the instinct model based on input information such as voice, image, and tactile sensation from the external environment to determine the motion. In other words, if the robot prepares an emotion model or an instinct model, it will be possible to realize realistic communication with humans at a higher intellectual level. In order for a robot to perform autonomous movements in response to changes in the environment, in the past, actions were described by a combination of simple action descriptions that received the information from a certain observation result and took action. By mapping behaviors to these inputs, it is possible to express complex behaviors that are not unique by introducing functions such as randomness, internal state (emotion / instinct), learning, and growth. However, with such a simple action mapping method, if you lose sight of the ball even for a moment in an action such as chasing and kicking the ball, information about the ball is not retained, so you should take an action such as searching for the ball from the beginning. And the use of recognition results is inconsistent. In addition, the robot needs to solve the correspondence problem of which recognition result corresponds to which previously observed result when a plurality of balls are observed. In addition, in order to utilize more complicated recognition results such as face recognition, skin color recognition, and voice recognition, which skin color area corresponds to which person on the face, and which person's voice is this voice. It is necessary to solve the problem. [Disclosure of Invention] An object of the present invention is to provide an excellent autonomous robot device having a function of recognizing information in the outside world such as an image or sound and reflecting its own behavior on the information, and a control method thereof. A further object of the present invention is to provide an excellent robot device and its control method capable of integrating a plurality of recognition results from the outside world, treating them as meaningful symbol information, and performing advanced behavior control. is there. A further object of the present invention is to utilize a more complicated recognition result such as a problem of correspondence with a previously observed recognition result to determine which skin color region corresponds to which person on the face and which person this voice corresponds to. It is an object of the present invention to provide an excellent robot device capable of solving voice or the like and a control method thereof. The present invention has been made in consideration of the above problems, and is a robot device capable of autonomous operation in response to changes in the external environment or a control method thereof. One or more recognition means or recognition steps that recognize the external environment, A target detection means or step that detects an object existing in the external environment based on the recognition result by the recognition means or step. A target storage means or step that holds target information about an object detected by the target detecting means or step. A robot device or a control method thereof. Here, the recognition means or the recognition step can recognize a voice, a user's utterance content, an object color, a user's face, and the like in an external environment, for example. The robot device or its control method according to the present invention may further include an action control means or step for controlling an action executed by the robot device itself based on the target information stored in the target storage means or step. Good. Here, the target storage means or step holds the recognition result of the object by the recognition means or step, the position of the object, and the posture information as the target information. At this time, the recognition result recognized in the coordinate system of the recognition means or step is converted into a fixed coordinate system based on the body of the robot device to hold the position and posture information of the object. It is possible to integrate information on multiple cognitive coordinates. For example, even if the robot moves its neck or the like to change the posture of the sensor, the position of the object as seen from the upper module (application) that controls the behavior of the robot based on the target information remains the same. .. Further, a target deletion means or step for deleting target information about an object that has not been recognized for a predetermined time or longer by the recognition means or step may be further provided. Further, it is determined whether the target not recognized by the recognition means or the step is within the recognition range or outside the range, and if it is outside the range, the corresponding target information of the target storage means is retained, and the target is within the range. For example, a forgetting means for forgetting the corresponding target information of the target storage means may be further provided. Further, the target associative means or step for finding the relationship between the target information may be further provided by using the same or overlapping recognized results as the same information about the object. Further, the target associative means or the step for finding the relationship between the target information may be further provided, with the results recognized by the two or more recognition means having the same position or overlapping as information on the same object. Further, the target storage means or step may add information recognized later with respect to the same object and hold it as one target information. Further, the robot device or its control method according to the present invention is detected by the event detecting means or step for detecting an event generated in the external environment based on the recognition result by the recognition means or step, and the event detecting means or step. It may further include event storage means or steps that retain information about the events in chronological order of occurrence. Here, the event detecting means or step may handle information related to changes in the external situation such as the recognized utterance content, the appearance and / or disappearance of an object, and the action executed by the robot device itself as an event. Further, an action control means or step for controlling an action executed by the robot device itself may be further provided based on the stored event information. The robot device or its control method according to the present invention can integrate a plurality of recognition results from the outside world and treat them as meaningful symbol information to perform advanced behavior control. In addition, since the individual recognition results notified asynchronously can be integrated and then passed to the action module, the action module facilitates the handling of information. Therefore, by using more complicated recognition results such as the problem of correspondence with the previously observed recognition results, it is possible to determine which skin color region corresponds to which person on the face, which person's voice is this voice, and so on. Can be solved. Further, according to the robot device according to the present invention or the control method thereof, since the information on the recognized observation result is stored as a memory, the observation result may not come temporarily during the period of autonomous action. However, it seems that an object is always perceived by a higher-level module such as an application that controls the behavior of the aircraft. As a result, it becomes resistant to mistakes in the recognizer and noise of the sensor, and a stable system that does not depend on the notification timing of the recognizer can be realized. Further, according to the robot device or the control method thereof according to the present invention, since the related recognition results are linked, it is possible to make an action judgment using the related information in a higher-level module such as an application. For example, the robot apparatus, based on the call is voice, it is possible to draw out the name of the person, it is possible to answer such as "Hello, XXX-san." The response of the greeting. Further, according to the robot device or the control method thereof according to the present invention, the information outside the field of view of the sensor is immediately retained without being forgotten, so that even if the robot loses sight of the object once, it can be searched for later. it can. Further, according to the robot device or the control method thereof according to the present invention, even if the information is insufficient from the viewpoint of the recognizer alone, other recognition results may be supplemented, so that the recognition performance of the entire system can be improved. improves. Still other objectives, features and advantages of the present invention will be clarified by more detailed description based on the embodiments of the present invention and the accompanying drawings described below. [Best mode for carrying out the invention] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.<u style="single">A. Configuration of leg-type mobile robot</u>1 and 2 show a view of the leg-type mobile robot 100 used for carrying out the present invention standing upright from both the front and the rear. This leg-type mobile robot 100 is a type called "humanoid" or "humanoid", and as will be described later, it is possible to autonomously control behavior based on the recognition result of an external stimulus such as voice or image. it can. As shown in the figure, the leg-type mobile robot 100 is composed of two left and right lower limbs that perform leg-type movement, a trunk, left and right upper limbs, and a head. Each of the left and right lower limbs is composed of a thigh, a knee joint, a shin, an ankle, and a palm, and is connected by a hip joint at substantially the lowermost end of the trunk. The left and right upper limbs are composed of an upper arm, an elbow joint, and a forearm, and are connected by shoulder joints at the left and right lateral edges above the trunk. In addition, the head is connected to the center of the uppermost end of the trunk by a neck joint. Inside the trunk unit, control units that are not visible in FIGS. 1 and 2 are deployed. This control unit is equipped with a controller (main control unit) that processes drive control of each joint actuator that constitutes the leg-type mobile robot 100, external inputs from each sensor (described later), a power supply circuit, and other peripheral devices. It is a housing. The control unit may also include a communication interface and a communication device for remote control. FIG. 3 schematically shows a joint degree of freedom configuration included in the leg-type mobile robot 100 according to this embodiment. As shown in the figure, the leg-type mobile robot 100 includes an upper body including two arms and a head 1, a lower limb consisting of two legs that realizes a moving motion, and a trunk portion that connects the upper limbs and the lower limbs. It is composed of and. The neck joint that supports the head 1 has three degrees of freedom: the neck joint yaw axis 2, the neck joint pitch axis 3, and the neck joint roll axis 4. In addition, each arm includes a shoulder joint pitch axis 8, a shoulder joint roll axis 9, an upper arm yaw axis 10, an elbow joint pitch axis 11, a forearm yaw axis 12, a wrist joint pitch axis 13, and a wrist joint roll. It is composed of a shaft 14 and a hand part 15. The hand 15 is actually an articulated, multi-degree-of-freedom structure containing multiple fingers. However, since the movement of the hand 15 itself has little contribution or influence on the posture stability control and walking movement control of the robot 100, it is assumed in this specification to have zero degrees of freedom. Therefore, it is assumed that each of the left and right arms has 7 degrees of freedom. Further, the trunk has three degrees of freedom: a trunk pitch axis 5, a trunk roll axis 6, and a trunk yaw axis 7. The left and right legs that make up the lower limbs include a hip joint yaw axis 16, a hip joint pitch axis 17, a hip joint roll axis 18, a knee joint pitch axis 19, an ankle joint pitch axis 20, and a joint roll axis 21. , Foot (sole or flat) 22. The foot (sole) 22 of the human body is actually a structure including a sole with multiple joints and multiple degrees of freedom, but the sole of the leg-type mobile robot 100 according to this embodiment has zero degrees of freedom. To do. Therefore, each of the left and right legs is composed of 6 degrees of freedom. Summarizing the above, the leg-type mobile robot 100 according to this embodiment has a total of 3 + 7 × 2 + 3 + 6 × 2 = 32 degrees of freedom. However, the leg-type mobile robot 100 is not necessarily limited to 32 degrees of freedom. Needless to say, the degree of freedom, that is, the number of joints, can be increased or decreased as appropriate according to design / manufacturing constraints and required specifications. The above-mentioned degrees of freedom of each joint of the leg-type mobile robot 100 are actually realized as active movements by the actuator. Due to various demands such as eliminating extra bulges on the appearance of the device to approximate the natural shape of a human being and controlling the attitude of an unstable structure such as bipedal walking, the joint actuator is compact and compact. It is preferably lightweight. In the present embodiment, it is decided to mount a small AC servo actuator of a type directly connected to the gear and having the servo control system integrated into one chip and built in the motor unit. The small AC servo actuator applicable to the leg-type mobile robot 100 is disclosed in, for example, Japanese Patent Application Laid-Open No. 2000-299970 (Japanese Patent Application No. 11-33386) which has already been assigned to the applicant. ing.<u style="single">B. Configuration of robot behavior control system</u>The leg-type mobile robot 100 according to the present embodiment can perform behavior control according to a recognition result of an external stimulus and a change in an internal state. FIG. 4 schematically shows the basic architecture of the behavior control system 50 adopted in the leg-type mobile robot 100 according to the present embodiment. Object-oriented programming can be incorporated into the illustrated behavioral control system 50. In this case, each software is handled in module units called "objects" that integrate data and processing procedures for the data. In addition, each object can exchange data and invoke by the inter-object communication method using message communication and shared memory. The behavior control system 50 includes a visual recognition function unit 51, an auditory recognition function unit 52, and a contact recognition function unit 53 in order to recognize the external environment (Environments). The visual recognition function unit (Video) 51 is, for example, a CCD (Charge Coupled). Device: Charge-coupled device) Based on the captured image input via an image input device such as a camera, image recognition processing such as face recognition and color recognition and feature extraction are performed. The visual recognition function unit 51 is composed of a plurality of objects such as "MultiColorTracker", "FaceDetector", and "FaceIdentify", which will be described later. The audio recognition function unit (Audio) 52 voice-recognizes voice data input via a voice input device such as a microphone, extracts features, and recognizes a word set (text). The auditory recognition function unit 52 is composed of a plurality of objects such as "AudioRecog" and "AuthurDecoder" which will be described later. The contact recognition function unit (Tactile) 53 recognizes a sensor signal from a contact sensor built in, for example, the head of the aircraft, and recognizes an external stimulus such as "struck" or "struck". Internal Status Management Department (ISM) The Manager) 54 is equipped with an instinct model and an emotion model, and the robot 100 responds to an external stimulus (ES: External Stimula) recognized by the above-mentioned visual recognition function unit 51, auditory recognition function unit 52, and contact recognition function unit 53. Manage internal states such as instincts and emotions. The emotion model and the instinct model have recognition results and action histories as inputs, and manage emotion values and instinct values, respectively. The behavior model can refer to these emotional values and instinct values. The short term memory 55 is a functional module that holds targets and events recognized from the external environment by the above-mentioned visual recognition function unit 51, auditory recognition function unit 52, and contact recognition function unit 53 for a short period of time. For example, the input image from the camera is stored for a short period of about 15 seconds. The long term memory 56 is used to retain information obtained by learning, such as the name of an object, for a very long period of time. The long-term memory unit 56, for example, associates and memorizes changes in the internal state from an external stimulus in a certain action module. However, since the associative memory in the long-term memory unit 56 is not directly related to the gist of the present invention, the description thereof is omitted here. The behavior control of the leg-type mobile robot 100 according to the present embodiment is performed by the "reflex behavior" realized by the reflex behavior unit 59, the "business condition-dependent behavior" realized by the situation-dependent behavior hierarchy 58, and the contemplation behavior hierarchy 57. It is roughly divided into "contemplation behavior" that is realized. The deliberative layer 57 performs a relatively long-term action plan of the leg-type mobile robot 100 based on the memory contents of the short-term memory unit 55 and the long-term memory unit 56. A pondering action is an action that is performed by making an inference or a plan to realize it according to a given situation or a command from a human being. Such reasoning and planning may require more processing time and computational load than the reaction time for the robot 100 to maintain interaction, so reflex behaviors and situation-dependent behaviors respond in real time while pondering behaviors. Make inferences and plans. In the Situated Behaviors Layer 58, the leg-type mobile robot 100 is currently placed based on the stored contents of the short-term memory unit 55 and the long-term memory unit 56 and the internal state managed by the internal state management unit 54. Control behaviors that are responsive to the situation. The situation-dependent action hierarchy 58 prepares a state machine for each action, classifies the recognition result of external information input by the sensor according to the action or situation before that, and expresses the action on the aircraft. To do. In addition, the situation-dependent action hierarchy 58 also realizes actions for keeping the internal state within a certain range (also called homeostasis action), and when the internal state exceeds the specified range, the internal state is changed to the corresponding range. Activate the action to make it easier to return to the range (actually, the action is selected in consideration of both the internal state and the external environment). Situation-dependent behaviors have a slower reaction time than reflex behaviors (discussed below). The contemplation action hierarchy 57 and the situation-dependent action hierarchy 58 can be implemented as applications. The Reflex Actions (ConfigurationDependentActionsAndReactions) 59 is a functional module that realizes reflexive aircraft movements in response to external stimuli recognized by the above-mentioned visual recognition function section 51, auditory recognition function section 52, and contact recognition function section 53. Is. The reflex behavior is basically an behavior that directly receives the recognition result of the external information input by the sensor, classifies it, and directly determines the output behavior. For example, behaviors such as chasing or nodding a human face are preferably implemented as reflex behaviors. In the leg-type mobile robot 100 according to the present embodiment, the short-term memory unit 55 temporally and spatially obtains the results of a plurality of recognizers such as the above-mentioned visual recognition function unit 51, auditory recognition function unit 52, and contact recognition function unit 53. It is integrated to maintain consistency and provides the perception of each object in the external environment as short-term memory to behavior control modules such as the Situation-Dependent Behavioral Hierarchy (SBL) 58. Therefore, on the behavior control module side configured as a higher-level module, it is possible to integrate a plurality of recognition results from the outside world and treat them as meaningful symbol information to perform advanced behavior control. In addition, by using more complicated recognition results such as the problem of correspondence with the previously observed recognition results, it is possible to determine which skin color region corresponds to which person on the face and which person's voice is this voice. Can be solved. In addition, since the short-term memory unit 55 holds information about the recognized observation result as a memory, the behavior of the aircraft is controlled even if the observation result does not come temporarily during the period of autonomous action. It is possible to make an object always appear to be perceived by a higher-level module such as an application. For example, information outside the field of view of the sensor is immediately remembered, so that even if the robot loses sight of an object, it can be found later. As a result, it becomes resistant to mistakes in the recognizer and noise of the sensor, and a stable system that does not depend on the notification timing of the recognizer can be realized. Further, even if the information is insufficient from the viewpoint of the recognizer alone, other recognition results may be supplemented, so that the recognition performance of the entire system is improved. In addition, since related recognition results are linked, it is possible to make action decisions using related information in higher-level modules such as applications. For example, a robot device can derive the name of a person based on the called voice. As a result, it is possible to notes, such as answer such as "Hello, XXX-san." The response of the greeting. FIG. 5 shows the flow of actions by each object constituting the action control system 50 shown in FIG. In the figure, circled objects are entities called "objects" or "processes". The entire system operates by communicating asynchronously with each other. Each object exchanges data and invokes by message communication and inter-object communication method using shared memory. The functions of each object will be described below.<u style="single">AudioRecog:</u>It is an object that receives voice data from a voice input device such as a microphone and performs feature extraction and voice section detection. Further, when the microphone is stereo, the sound source direction can be estimated in the horizontal direction. If it is determined to be an audio section, the feature amount and sound source direction of the audio data in that section are sent to Arther Decoder (described later).<u style="single">ArthurDecoder:</u>It is an object that performs voice recognition using the voice features received from AudioRecog and the voice dictionary and syntax dictionary. The set of recognized words is sent to the Short Term Memory 55.<u style="single">MultiColorTracker:</u>It is an object that performs color recognition, receives image data from an image input device such as a camera, extracts a color area based on a plurality of color models that it has in advance, and divides it into continuous areas. Information such as the position, size, and feature amount of each divided area is output and sent to the short term memory (ShortTermMemory) 55.<u style="single">FaceDetector:</u>It is an object that detects the face area from the image frame, receives image data from an image input device such as a camera, and reduces and converts it into a 9-step scale image. A rectangular area corresponding to a face is searched from all these images. It reduces the overlapping candidate areas, outputs information such as the position, size, and feature amount of the area finally determined to be the face, and sends it to Face Identify (described later).<u style="single">FaceIdentify:</u>It is an object that identifies the detected face image, receives a rectangular area image showing the face area from FaceDetector, compares which person in the person dictionary on hand corresponds to the person, and identifies the person. Do it. In this case, the face image is received from the face detection, and the ID information of the person is output together with the position and size information of the face image area as output.<u style="single">ShortTerm Memory:</u>It is an object that holds information about the external environment of the robot 100 for a relatively short time, receives the voice recognition result (word, sound source direction, certainty) from ArthurDecoder, and the position, size and face area of the skin color area from MultiColorTracker. , Receives the size and receives the person's ID information etc. from FaceIdentify. In addition, the direction of the robot's neck (joint angle) is received from each sensor on the body of the robot 100. Then, by using these recognition results and sensor outputs in an integrated manner, information on who is currently, what person the spoken words belong to, and what kind of dialogue has been conducted with that person so far is obtained. save. Physical information about such an object, that is, a target and an event (history) viewed in the time direction are output and passed to a higher module such as a situation-dependent behavior hierarchy (SBL).<u style="single">LocatedBehaviorLayer:</u>It is an object that determines the behavior (action depending on the situation) of the robot 100 based on the information from the short term memory (short-term memory) described above. You can evaluate and perform multiple actions at the same time. You can also switch actions to put the aircraft to sleep and activate another action.<u style="single">ResourceManager:</u>It is an object that arbitrates the resources of each hardware of the robot 100 in response to the output command. In the example shown in FIG. 5, resource arbitration is performed between the object that controls the speaker for voice output and the object that controls the motion of the neck.<u style="single">SoundPerformerTTS:</u>It is an object for voice output, and voice synthesis is performed according to the text command given by SituatedBehaviorLayer via ResourceManager, and voice output is performed from the speaker on the robot 100.<u style="single">HeadMotionGenerator:</u>It is an object that calculates the joint angle of the neck in response to receiving a command to move the neck from SituatedBehaviorLayer via ResourceManager. When the "track" command is received, the joint angle of the neck facing the direction in which the object exists is calculated and output based on the position information of the object received from ShortTermMemory.<u style="single">C. Short-term memory</u>In the leg-type mobile robot 100 according to the present embodiment, the ShortTermMemory 55 integrates the results of a plurality of recognizers related to external stimuli so as to maintain consistency in time and space, and has meaning. It is designed to be treated as symbol information. As a result, higher-level modules such as the Situation-Dependent Behavioral Hierarchy (SBL) 58 utilize more complex recognition results, such as correspondence problems with previously observed recognition results, to accommodate which skin-colored area corresponds to which person on the face. It is possible to understand whether or not this voice is the voice of which person. The short-term storage unit 55 is composed of two types of memory objects, a target memory and an event memory. The target memory integrates the information from each recognition function unit 51 to 53 and holds the information about the currently perceived object, that is, the target. Therefore, when the target object disappears or appears, the corresponding target is deleted from the storage area (Garbage Collector) or newly generated. In addition, one target can be represented by multiple recognition attributes (Target Associate). For example, it is an object (human face) that emits a voice with a skin color and a face pattern. The position and orientation information of the object (target) held in the target memory is not the sensor coordinate system used in each recognition function unit 51 to 53, but a specific part on the machine such as the trunk of the robot 100. It is expressed in the world coordinate system fixed in a predetermined place. Therefore, the short-term memory unit (STM) 55 constantly monitors the current value (sensor output) of each joint of the robot 100 and converts the sensor coordinate system to this fixed coordinate system. This makes it possible to integrate the information of each recognition function unit 51 to 53. For example, even if the robot 100 moves its neck or the like to change the posture of the sensor, the position of the object as seen from the behavior control module such as the situation-dependent behavior hierarchy (SBL) remains the same, making it easy to handle the target. Become. The event memory is an object that stores events from the past to the present that have occurred in the external environment in chronological order. Events handled in the event memory include information on changes in external conditions such as the appearance and disappearance of targets, speech-recognized words, and changes in one's own behavior and posture. The event contains a change of state for a target. Therefore, by including the ID of the corresponding target as the event information, it is possible to search the above-mentioned target memory for more detailed information about the event that has occurred. One point of the present invention is to detect a target by integrating two or more types of sensor information having different attributes. For the distance calculation between the sensor result and the storage result, for example, a normal distance calculation (polar coordinates are angles) is used, which deals with the distance obtained by subtracting the size of both by half from the distance at the center position of the target. When integrating the sensor information about voice and the sensor information about face, if the target in the short-term memory contains the recognition result of face when the recognition result about voice is obtained, the normal distance Integrate based on calculations. Otherwise, the distance is treated as infinite. When integrating sensor information related to color, if the color recognition result and the stored target color are the same color, the normal distance x 0.8 is set as the distance, and in other cases, the normal distance x 4.0 is set as the distance. To do. In addition, when integrating the sensor result and the storage result, the average of the elevation angle and the horizontal angle should be within 25 degrees in the case of polar coordinate system representation, and the normal distance should be within 50 cm in the case of Cartesian coordinate system representation. Usually an integration rule. In addition, the difference due to each recognition result can be dealt with by adding a weight to the calculation method of the distance between the recognition result and the target. FIG. 6 schematically shows how the short-term memory unit 55 operates. In the example shown in the figure, the operation when the face recognition result (FACE), the voice recognition, and the recognition result (VOICE) of the sound source direction are processed at different timings and notified to the short-term memory unit 55. It is represented (however, it is drawn in a polar coordinate system with the body of the robot 100 as the origin). In this case, since each recognition result is temporally and spatially close (overlapping), it is judged that it is one object having facial attributes and voice attributes, and the target memory is updated. are doing. 7 and 8 show the flow of information entering the target memory and the event memory in the short-term storage unit 55 based on the recognition results in the recognition function units 51 to 53, respectively. As shown in FIG. 7, a target detector for detecting a target from the external environment is provided in the short-term storage unit 55 (STM object). This target detector adds a new target or reflects an existing target in the recognition result based on the recognition results by each recognition function unit 51 to 53 such as voice recognition result, face recognition result, and color recognition result. Update to. The detected target is kept in the target memory. In addition, the target memory has functions such as a Garbage Collector that searches for and erases targets that are no longer observed, and a Target Associate that determines the relevance of multiple targets and associates them with the same target. There is. The garbage collector is realized by decrementing the certainty of the target over time and deleting the target whose certainty is below the predetermined value. In addition, the target associate can identify the same target by having spatial and temporal closeness between targets having the same attribute (recognition type) and similar features. The above-mentioned situation-dependent behavioral hierarchy (SBL) is an object that becomes a client (STM client) of the short-term storage unit 55, and periodically receives a notification (Notify) of information about each target from the target memory. In this embodiment, the STM proxy class copies the target to a client-local work area independent of the short-term memory 55 (STM object) and always keeps the latest information. Then, the desired target is read from the local target list (Target of Interest) to determine the schema, that is, the action module. Further, as shown in FIG. 8, an event detector for detecting an event occurring in the external environment is provided in the short-term storage unit 55 (STM object). This event detector detects the creation of a target by the target detector and the deletion of the target by the garbage collector as an event. Further, when the recognition result by the recognition function units 51 to 53 is voice recognition, the utterance content becomes an event. The events that occur are stored as an event list in the event memory in the order of the time they occurred. The above-mentioned situation-dependent action hierarchy (SBL) is an object that becomes a client (STM client) of the short-term storage unit 55, and receives an event notification (Notify) from the event memory every moment. In this embodiment, the STM proxy class copies the event list to a client-local work area independent of the short-term memory 55 (STM object). It then reads the desired event from the local event list to determine the schema, or action module. The executed action module is detected by the event detector as a new event. In addition, old events are sequentially discarded from the event list in the form of, for example, FIFO (Fast In Fast Out). FIG. 9 shows the processing operation of the target detector in the form of a flowchart. Hereinafter, the process of updating the target memory by the target detector will be described with reference to this flowchart. When the recognition results from each recognition function unit 51 to 53 are received (step S1), the joint angle data at the same time as the recognition result is searched, and based on that, each recognition result is obtained from the sensor coordinate system to the world fixed coordinate system. Perform coordinate conversion to (step S2). Next, one target is taken out from the target memory (Target of Interest) (step S3), and the position and time information of the target is compared with the position and time information of the recognition result (step S4). If the positions overlap and the measurement times are close, it is judged that the target and the recognition result match. If they match, it is further checked whether there is information of the same recognition type as the recognition result in the target (step S5). Then, if there is information of the same recognition type, it is further checked whether or not the features match (step S6). If the features match, the recognition result of this time is reflected (step S7), the target position and observation time of the target are updated (step S8), and the entire processing routine is terminated. On the other hand, if the features do not match, a new target is generated, the current recognition result is assigned to it (step S11), and the entire processing routine is terminated. If it is determined in step S5 that there is no information of the same recognition type in the targets with matching recognition information, the recognition result is added for the target (step S9), and the target position and observation time are set. Update (step S9) to perform the entire processing routine. If it is determined in step S4 that the extracted target does not match the position and time information of the recognition result, the next target is sequentially extracted (step S10), and the same processing as described above is repeated. If a target that finally matches the recognition result cannot be found, a new target is generated, the current recognition result is assigned to it (step S11), and the entire processing routine is terminated. As already mentioned, the addition or update of a target by the target detector is an event. FIG. 10 shows the processing procedure for the garbage collector to delete the target from the target memory in the form of a flowchart. The garbage collector is called and activated on a regular basis. First, the measurement range of the sensor is converted to the world fixed coordinate system (step S21). Then, one target is fetched from the target memory (Target of Interest) (step S22). Then, it is checked whether or not the extracted target is within the measurement range of the sensor (step S23). Once the target is within the measurement range of the sensor, decrement its confidence over time (step S24). If there is a possibility that the target does not exist because the target information has not been updated even though it is within the measurement range of the sensor, the certainty is lowered. Information is retained for targets outside the measurement range. Then, when the certainty of the target falls below the predetermined threshold TH (step S25), the target is deleted from the target memory as if it is no longer observed (step S26). The process of updating the certainty and deleting the target as described above is repeated for all the targets (step S27). As already mentioned, the deletion of the target by the garbage collector is an event. FIG. 11 shows the data representation of the target memory. As shown in the figure, a target is a list having a plurality of recognition results called Associated Target as elements. For this reason, the target can have any number of related recognition results. At the top of the list, physical information such as the position and size of this target, which is comprehensively judged from all recognition results, is stored, and recognition results such as voice, color, and face are stored. Follows this. AssociatedTargets are expanded in shared memory and are created or erased in memory as the target appears or disappears. By referring to this memory, the upper module can obtain information on the object currently under the consciousness of the robot 100. In addition, FIG. 12 shows the data representation of the event memory. As shown in the figure, each event is represented by a structure called STMEvent, which consists of a data field common to all events and a recognizer-specific field that detects the event. The event memory is a storage area that is included as an element of the list of these STMEvent structures. The upper module can search for the desired event by using the corresponding data from the data in the data field of STMEvent. When searching the history of a dialogue with a specific person by such an event search, all the words spoken by that person are searched using the target ID corresponding to that person and the event type called SPEECH (speech). It becomes possible to list. Further, FIG. 13 shows an example of a structure for storing the recognition results in the recognition function units 51 to 53. The data fields of the structure shown in the figure are provided with data fields such as feature quantities that depend on the respective recognition function units 51 to 53 and physical parameter fields such as position and size velocity that do not depend on recognition. There is.<u style="single">D. Robot-based dialogue processing</u>In the leg-type mobile robot 100 according to the present embodiment, the results of a plurality of recognizers related to external stimuli are integrated so as to maintain consistency in time and space, and are treated as meaningful symbol information. There is. This makes it possible to utilize more complicated recognition results such as the problem of correspondence with previously observed recognition results, such as which skin color region corresponds to which person on the face, and which person's voice is this voice. It is possible to solve. Hereinafter, the interactive processing between the robot 100 and the users A and B by the robot 100 will be described with reference to FIGS. 14 to 16. First, as shown in FIG. 14, when user A calls "Masahiro (robot name) -kun!", Sound direction detection, voice recognition, and face identification are performed by each recognition function unit 51 to 53, and the call is called. Situation-dependent actions such as tracking the face of user A and initiating a dialogue with user A are performed. Next, as shown in FIG. 15, when the user B calls "Masahiro (robot name) -kun!", Each recognition function unit 51 to 53 performs sound direction detection, voice recognition, and face identification. Situation-dependent behaviors such as interrupting a conversation with user A (but preserving the context of the conversation) and then turning to the called direction to track user B's face or start a conversation with user B. Is performed. Then, as shown in FIG. 16, user A shouts "Hey!" To urge the conversation to continue, and this time after interrupting the conversation with user B (but preserving the context of the conversation). Situation-dependent actions are taken to turn to the called direction, track User A's face, and resume dialogue with User A based on the stored context.<u style="single">Supplement</u>The present invention has been described in detail with reference to the specific examples. However, it is self-evident that a person skilled in the art can modify or substitute the embodiment without departing from the gist of the present invention. The gist of the present invention is not necessarily limited to a product called a "robot". That is, as long as it is a mechanical device that performs a motion similar to a human movement by using an electric or magnetic action, even if it is a product belonging to another industrial field such as a toy, the present invention is similarly applied. Can be applied. In short, the present invention has been disclosed in the form of an example, and the contents of the present specification should not be construed in a limited manner. In order to judge the gist of the present invention, the column of claims described at the beginning should be taken into consideration. [Industrial applicability] According to the present invention, it is possible to provide an excellent autonomous robot device having a function of recognizing information in the outside world such as an image or sound and reflecting its own behavior on the information, and a control method thereof. Further, according to the present invention, it is provided an excellent robot device capable of integrating a plurality of recognition results from the outside world, treating them as meaningful symbol information, and performing advanced behavior control, and a control method thereof. Can be done. Further, according to the present invention, by utilizing a more complicated recognition result such as a problem of correspondence with a previously observed recognition result, which skin color region corresponds to which person on the face, and which person this voice corresponds to. It is possible to provide an excellent robot device and a control method thereof that can solve the voice of the person. Since the robot device according to the present invention can integrate the individual recognition results notified asynchronously and then pass them to the action module, the action module facilitates the handling of information. Further, since the robot device according to the present invention holds information on the recognized observation result as a memory, the behavior of the aircraft even if the observation result does not come temporarily during the period of autonomous action. It always appears that an object is perceived by a higher-level module such as a controlling application. As a result, it becomes resistant to mistakes in the recognizer and noise of the sensor, and a stable system that does not depend on the notification timing of the recognizer can be realized. Further, in the robot device according to the present invention, since the related recognition results are linked, it is possible to make an action judgment using the related information in a higher-level module such as an application. For example, the robot apparatus, based on the call is voice, it is possible to draw out the name of the person, it is possible to answer such as "Hello, XXX-san." The response of the greeting. Further, since the robot device according to the present invention immediately retains information outside the field of view of the sensor, even if the robot loses sight of the object once, it can be searched for later. Further, in the robot device according to the present invention, even if the information is insufficient when viewed from the recognizer alone, other recognition results may be supplemented, so that the recognition performance of the entire system is improved. [Simple explanation of drawings] FIG. 1 is a view showing a state in which the leg-type mobile robot 100 used for carrying out the present invention is viewed from the front. FIG. 2 is a view showing a rear view of the leg-type mobile robot 100 used for carrying out the present invention. FIG. 3 is a diagram schematically showing a degree of freedom configuration model included in the leg-type mobile robot 100 according to the embodiment of the present invention. FIG. 4 is a diagram schematically showing the architecture of the behavior control system 50 adopted in the leg-type mobile robot 100 according to the embodiment of the present invention. FIG. 5 is a diagram schematically showing the flow of movement in the behavior control system 50. FIG. 6 is a diagram schematically showing how the short-term memory unit 55 operates. FIG. 7 is a diagram showing a flow of information entering the target memory based on the recognition results of the recognition function units 51 to 53. FIG. 8 is a diagram showing a flow of information entering the event memory based on the recognition results of the recognition function units 51 to 53. FIG. 9 is a flowchart showing the processing operation of the target detector. FIG. 10 is a flowchart showing a processing procedure for the garbage collector to delete the target from the target memory. FIG. 11 is a diagram showing the data representation of the target memory. FIG. 12 is a diagram showing a data representation of the event memory. FIG. 13 is a diagram showing an example of a structure for storing the recognition results in the recognition function units 51 to 53. FIG. 14 is a diagram showing an example of interactive processing between the robot 100 and the users A and B. FIG. 15 is a diagram showing an example of interactive processing between the robot 100 and the users A and B. FIG. 16 is a diagram showing an example of interactive processing between the robot 100 and the users A and B.
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US10583559B2 | Cited by | United States of America | Applicant |
| JP2017514709A | Cited by | Japan | Search report |
| KR20170027704A | Cited by | Republic of Korea | Search report |
| KR20170029409A | Cited by | Republic of Korea | Search report |
| KR101273300B1 | Cited by | Republic of Korea | Search report |
| JP2017514709A | Cited by | Japan | Search report |
| JP2017513724A | Cited by | Japan | Search report |
| JP2017513724A | Cited by | Japan | Search report |
| JP2001188555A | Cites | Japan | – |
| JP09081205A | Cites | Japan | – |
| JP2001157981A | Cites | Japan | – |
11 members in 6 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 2001323259 | Japan | A | |
| 2001323259 | Japan | A | |
| 2001323259 | Japan | – | |
| 0210921 | Japan | W | |
| 0210921 | Japan | W | |
| 20012001323259 | – | – | – |
| 2002010921 | – | – | – |
| JP20010323259 | – | – | – |
| WO2002JP10921 | – | – | – |
Members11
| Document | Office | Kind | |
|---|---|---|---|
| WO03035334A1 | World Intellectual Property Organization (WIPO) | A1 | |
| CN1487871A | China | A | |
| KR20040050055A | Republic of Korea | A | |
| US2004117063A1 | United States of America | A1 | |
| EP1439039A1 | European Patent Office (EPO) | A1 | |
| US6850818B2 | United States of America | B2 | |
| JPWO2003035334A1 | Japan | A1 | |
| CN1304177C | China | C | |
| EP1439039A4 | European Patent Office (EPO) | A4 | |
| KR100898435B1 | Republic of Korea | B1 | |
| JP4396273B2This record | Japan | B2 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4396273
- Publication, DOCDB
- 4396273
- Publication, EPODOC
- JP4396273B
- Application
- 2003537879
- Application, DOCDB
- 2003537879
- Application, EPODOC
- JP20030537879
Titles2
- Japanese
- ロボット装置及びその制御方法
- English
- Robot device and its control method
Classification
- CPC, 5
- G06N3/008
- B25J5/00
- G05B19/4097
- G05B2219/33051
- B25J13/00
- IPC, 5
- B25J13 00
- B25J5 00
- B25J13 08
- G05B19 4097
- G06N3 00