Video camera driving device
Abstract
(57) A summary and the purpose The video camera drive which can track a photographic subject automatically is offered without needing heat perception sensors, such as an infrared camera, etc. Composition If a part of image (for example, a part of a speaker's face) currently displayed on the speaker display monitor 5 is specified as an observation region, the moving vector of this domain will be computed as a domain characteristic vector, and the domain characteristic vector information S14 which shows this will be supplied to the phase control angle operation part 15. Thereby, the domain characteristic vector information S14 is changed into the angle vector information S15 which shows the horizontal and perpendicular angle of a camera fixed stand. Next, the camera fixed stand actuator 3 changes respectively the horizontal and perpendicular angle of a camera fixed stand according to the angle vector information S15. As a result, the angle of the video camera 4 follows a speaker's motion, and is controlled, and a speaker's image is displayed in the monitor's 5 center.
Term
No projected expiry on record.
- Priority and filed
- Published
- Today
2 claims: 1 independent, 1 dependent
- 1[Claims] 1. In a video camera driving device that drives a video camera according to the movement of the subject so that the subject is displayed at a predetermined position on a monitor. Area designation means for designating a predetermined area of the subject displayed on the monitor, and A motion vector calculation means for calculating a motion vector of a region designated by this area designation means, and a motion vector calculation means. A video characterized by including a camera driving means for driving the video camera so that the subject is displayed at a predetermined position on the monitor based on the motion vector calculated by the motion vector calculating means. Camera drive. 【特許請求の範囲】 【請求項1】 被写体がモニタ上の所定位置に表示されるように、この被写体の動きに合わせてビデオカメラを駆動するビデオカメラ駆動装置において、 前記モニタに表示された前記被写体の所定の領域を指定する領域指定手段と、 この領域指定手段により指定された領域の動きベクトルを算出する動きベクトル算出手段と、 この動きベクトル算出手段により算出された動きベクトルに基づいて、前記被写体が前記モニタ上の所定位置に表示されるように、前記ビデオカメラを駆動するカメラ駆動手段とを具備したことを特徴とするビデオカメラ駆動装置。
106 paragraphs, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Industrial application field]
The present invention relates to, for example, a video camera driving device that drives a video camera in accordance with the movement of the speaker so that the speaker is always displayed at a predetermined position on a monitor in a video conferencing system.
【0002】
[Conventional technology]
In recent years, various systems, so-called video conferencing systems, have been developed in which a speaker is photographed by a video camera and the video is transmitted to the other party of the conference via a communication network. In this video conferencing system, it is desirable that the speaker always be displayed in a predetermined position on the monitor.
【0003】
In order to meet this demand, a device that drives a video camera according to the movement of the speaker is required. As this video camera drive device, a device as described in Japanese Patent Application Laid-Open No. 4-103285 has been conventionally developed. This device integrates an infrared camera and a video camera that shoot the same object, and controls the horizontal and vertical angles of both cameras based on the heat of the speaker captured by the infrared camera. In detail, it has the following functions.
【0004】
First, the heat generated by the speaker is captured by an infrared camera, the height of the heat is converted into a predetermined number of block-shaped characters, and this is compared with a standard character. This standard character is set assuming that the image of the speaker is, for example, the center of the television monitor. Then, the angle of the infrared camera is adjusted so that the character based on the output signal of the infrared camera approaches the standard character.
【0005】
[Problems to be Solved by the Invention]
According to such a configuration, the angle of the video camera integrated with the infrared camera is also corrected according to the movement of the speaker, and the image of the speaker is always displayed in the center of the television monitor.
【0006】
However, in such a configuration, since the movement of the subject must be detected by using an infrared camera, there is a problem that the device becomes large-scale and expensive.
【0007】
The present invention has been made to solve the above-mentioned problems, and an object of the present invention is to provide a video camera driving device that does not require an infrared camera and can be miniaturized and inexpensive.
【0008】
[Means for solving problems]
In order to solve the above problems, the present invention comprises an area designating means for designating a predetermined area of a subject displayed on a monitor, and a motion vector calculating means for calculating a motion vector of the area designated by the area designating means. Based on the motion vector calculated by the motion vector calculating means, a camera driving means for driving the video camera is provided so that the subject is displayed at a predetermined position on the monitor.
【0009】
[Action]
In the above configuration, first, a predetermined area of the subject displayed on the monitor is designated by the area designating means. When this designation is made, the motion vector calculation means sequentially calculates the motion vector in the designated region. Then, the video camera is driven based on this calculated output. As a result, the angle of the video camera is corrected according to the movement of the subject, and the subject is always displayed at a predetermined position on the monitor, for example, in the center.
【0010】
[Example]
A: Configuration of Examples Hereinafter, examples of the present invention will be described with reference to the drawings.
【0011】
FIG. 1 is a block diagram showing a configuration of an embodiment of the present invention. It should be noted that this example is an example in which the present invention is applied to a video conferencing system.
【0012】
In the figure, the camera movement request switch 1 is placed at a fixed position on the desk of each attendee of the conference, and when operated by the attendees, the camera movement request information S1 indicating the movement request of the video camera 4. Is sent to the control angle calculation unit 2 for camera movement request.
【0013】
When the camera movement request control angle calculation unit 2 receives the camera movement request signal S1, the camera movement request angle vector information S2 indicating the angle of the video camera 4 for photographing the vicinity of the camera movement request switch 1 is captured by the camera. Output to the fixed base drive unit 3.
【0014】
The camera fixing base drive unit 3 drives the camera fixing base (not shown), and the video camera 4 is attached to the camera fixing base. The camera fixing base driving unit 3 drives the camera fixing base and controls the angle of the video camera 4 based on the angle vector information S2 for requesting camera movement or the angle vector information S15 described later. The video camera 4 captures a subject and outputs the image signal to the speaker display monitor 5 and the coding unit 16.
【0015】
The coding unit 16 performs various processing on the image signal supplied from the video camera 4 and transfers it to the other party speaker of the conference via the network NT, and is a component of the portion surrounded by the broken line in the figure. have. First, the AD conversion unit 6 converts the image signal supplied from the video camera 4 into a digital signal, and supplies the image data obtained as a result to the current frame buffer 7. The current frame buffer 7 stores image data for one frame at the present time (hereinafter referred to as current screen data S7), and supplies the stored current screen data S7 to the motion vector calculation unit 9 and the DCT / quantization unit 10. To do. The motion vector calculation unit 9 is M × M based on the front screen playback data S8 output by the previous frame buffer 8 that stores the previous playback image data and the current screen data S7 output by the current frame buffer 7. The motion vector information S9 of the macroblock composed of pixels is calculated. The DCT / quantization unit 10 performs predetermined processing (details will be described later) based on the motion vector information S9, the current screen data S7, and the previous screen reproduction data S8, and outputs the output signal S10 to the entropy encoding unit 11 and the reverse. It is supplied to the quantization / IDCT unit 12. The inverse quantization / IDCT unit 12 performs the reverse processing of the DCT / quantization unit 10, and the processing result is the reproduced image data one screen before. The image data output by the inverse quantization / IDT unit 12 is stored in the previous frame buffer 8. The entropy coding unit 11 performs Huffman coding that compresses information by utilizing the statistical properties of the signal, and the output signal is output as the coded image data S11.
【0016】
Further, the attention area designation unit 13 designates the attention area of the speaker displayed on the speaker display monitor 5 according to the operation of the operator, and provides the attention area designation information S13 indicating the designated attention area. Output to the region-specific motion vector calculation unit 14. The region-specific motion vector calculation unit 14 creates region-specific motion vector information S14 based on the attention region designation information S13 and the motion vector information S9, and supplies the region-specific motion vector information S14 to the control angle calculation unit 15.
【0017】
Next, the decoding unit 17 demodulates the coded data supplied from the other party of the conference via the network NT, and outputs the demodulated data obtained as a result. This demodulated data is converted into an analog signal by the DA conversion unit 18 and supplied to the other party image display monitor 19. The other party image display monitor 19 displays the image of the other party in the conference based on the demodulated analog signal supplied from the DA conversion unit 18.
【0018】
B: Operation of the example Next, the operation of this embodiment according to the above-described configuration will be described.
【0019】
First, when a certain attendee presses the camera movement request switch 1 provided on the desktop, the camera movement request information S1 is transmitted to the camera movement request control angle calculation unit 2. The camera movement request control angle calculation unit 2 stores the location assigned to each camera movement request switch 1 in advance, and the horizontal angle and vertical of the camera fixing base are used to move the shooting location of the video camera to that position. Outputs the angle vector information S2 for camera movement request including angle information. The camera fixing base drive unit 3 changes the horizontal angle and the vertical angle of the camera fixing base according to the angle vector information S2 for requesting camera movement. As a result, the video camera 4 captures the image of the subject to be photographed, for example, the speaker who has pressed the camera movement request switch 1. The image taken by the video camera 4 is displayed on the spot by the speaker display monitor 5, and after the information is compressed by the coding unit 16, it is transferred to the other party via the network NT. At the other party, the transferred coded data S11 is decoded and reproduced.
【0020】
Next, the attention area designation unit 13 specifies the attention area for the image displayed on the speaker display monitor 5. This designation is performed by a pointing device such as a mouse provided in the attention area designation unit 13, and for example, as shown in FIG. 2, a part of the speaker's face is designated. In the figure, R is a region of interest and P is a pointer that specifies the region of interest R.
【0021】
The attention area designation information S13 output by the attention area designation unit 13 is created as information uniquely representing the attention area R. For example, when the region of interest R is specified as a rectangle, the information indicates the horizontal and vertical coordinates of one vertex and the length of each side in the horizontal and vertical directions.
【0022】
(Operation of Coding Unit 16) Next, the processing contents of the coding unit 16 will be described. The coding processing method of the coding unit 16 in this embodiment is the same as the coding method specified in CCITT Recommendation H.261.
【0023】
First, the image data output from the video camera 4 is converted into a digital signal by the AD conversion unit 6 and then stored in the current frame buffer 7. The motion vector calculation unit 9 uses the current screen data S7 and the previous screen playback data S8 to perform motion compensation prediction, which is an image compression technology, for each macro block consisting of M × M pixels (M is a positive integer). And detect the motion vector.
【0024】
The detection of this motion vector is performed as follows. First, as shown in FIG. 3, for each macroblock MB1 on the current frame screen, a search area SA surrounding the macroblock MB2 at the same position as the macroblock MB1 is set on the previous frame screen. Next, a macroblock MB3 having the same size of M × M pixels as the macroblock is cut out from this search area SA. Next, the similarity between the data of this block MB3 and the data of macroblock MB1 is calculated. The above processing is performed for all macroblocks MB3 that can be cut out from the search area SA, and the macroblock MB3 with the highest similarity is detected. Finally, the difference between the position of macroblock MB3 and the position of macroblock MB2, which have the highest degree of similarity, is calculated. This difference is the motion vector v of the macroblock MB1.
【0025】
Here, the direction of the motion vector v is from the macroblock MB2 to the maximum similar macroblock MB3 in the search area SA, which is opposite to the actual movement of the subject. The component of the motion vector v is defined so that "1" represents the distance of one pixel.
【0026】
As described above, the motion vector calculation unit 9 outputs the motion vector v of each macroblock MB1 on the current frame screen as the motion vector information S9. Then, the DCT / quantization unit 10 sets the M × M pixel area (macroblock MB3) of the previous frame playback screen corresponding to each macroblock MB1 of the current frame screen to each macroblock MB1 based on the motion vector information S9. After moving in the direction opposite to the motion vector v, the density difference value between the image data in this area and the macroblock MB1 of the current frame screen is obtained. Then, the DCT / quantization unit 10 either DCT (discrete cosine transform) / quantizes the difference value itself and outputs it according to the magnitude of this difference, or DCT / quantum the macroblock MB1 itself of the current frame screen. And output. The discrete cosine transform is a kind of orthogonal transform, and the quantization is to find the quotient divided by an integer called a quantization step, both of which are generally used as a basic technique of code compression. Further, the macroblock output as the above-mentioned density difference value is called an inter macroblock, and the macroblock output as the current screen itself is called an intra macroblock.
【0027】
Next, the entropy code unit 11 compresses the information of the signal S10 by using statistical properties, and outputs the output as the coded image data S11.
【0028】
Further, the inverse quantization / IDCT unit 12 performs processing such as inverse quantization and inverse DCT on the output signal S10 of the DCT / quantization unit 10. In this case, after processing according to the difference between the inter macroblock and the intra macroblock, the output is transferred to the previous frame buffer 8.
【0029】
(Operation of the region eigenvector calculation unit 14) Next, the operation of the region eigenvector calculation unit 14 will be described. In the coding method specified in H.261 of CCITT, it is essential to calculate the motion vector v for each macroblock MB1 as described above. Therefore, the region-specific motion vector calculation unit 14 calculates the motion vector V peculiar to the region of interest R by using the motion vector information S9 for each macroblock MB1 calculated by the motion vector calculation unit 9 in the coding unit 16. I try to do it.
【0030】
That is, the region-specific vector calculation unit 14 receives the attention region designation information S13 output by the attention region designation unit 13 and the motion vector information S9 of each macroblock MB1 output by the motion vector calculation unit 9, and receives the region-specific motion vector from the following equation. V is calculated and output as the motion vector information S14 peculiar to the region of interest.
【0031】
[Number 1]
<img file="JPH06339056A_D0001.tif" />Here, the molecule on the right side of Equation 1 is the motion vector vi of each macroblock containing the pixels included in the region of interest R (hereinafter referred to as the macroblock belonging to the region of interest) MB1 (i) (i = 1 to N). And the product of the number mi of the pixels included in the macroblock MB1 (i) and also included in the region R of interest, all the macroblocks MB1 (1) to MB1 (N) belonging to the region R of interest. ), And the sum of them is taken. Further, the denominator on the right side of the equation 1 is included in each macroblock MB1 (i) belonging to the region of interest R, and the number mi of the pixels included in the region of interest R is included in all the macroblocks MB1 (1) belonging to the region of interest R. It is the sum of 1) ~ MB1 (N).
【0032】
Therefore, the fraction on the right side of Equation 1 represents the average motion vector v of the macroblocks MB1 (1) to MB (N) belonging to the region of interest R. Further, the negative sign on the right side of Equation 1 reverses the direction of the motion vector, and as a result, the region-specific motion vector V obtained by Equation 1 becomes a vector indicating the actual motion of the region of interest R. That is, the object to be photographed in the region of interest R has moved by the region eigenvector V.
【0033】
Then, the region eigenvector information S14 indicating the region eigenvector V is supplied to the control angle calculation unit 15 and converted into the angle vector information S15 indicating the horizontal angle and the vertical angle of the camera fixing base. This angle vector information S15 is obtained as follows.
【0034】
First, let Φ be the rotation angle when the object reflected in the center on the monitor moves by one pixel on the monitor when the camera fixing base is horizontally rotated. Then, if the horizontal component of the region eigenvector V is m, the horizontal component of the angle vector information S15 is approximated by mΦ. The vertical angle component is calculated in the same way.
【0035】
Next, the camera fixing base drive unit 3 changes the horizontal and vertical angles of the camera fixing base by the horizontal angle component and the vertical angle component of the angle vector information S15, respectively. As a result, the shooting angle of the video camera 4 is corrected according to the movement of the speaker, and the image displayed in the attention area R is always approximately in the center of the monitors (other party image display monitor 19 and speaker display monitor 5). Positioned.
【0036】
When the camera movement request switch 1 is newly pressed and the camera movement request angle vector information S2 is output in response to this, the camera fixing base drive unit 3 gives priority to the camera movement request angle vector information S2. , Drive the camera fixing base accordingly. Therefore, even if the control is performed by the angle vector information S15, when the camera movement request switch 1 is pressed, the camera fixing base is driven so as to be the shooting position stored in advance for the camera movement request switch 1. To. After that, the angle vector information S15 is ignored until the attention area R is set again by the attention area designation unit 13, and when the attention area R is set, the control by the angle vector information S15 is performed again.
【0037】
According to this embodiment described in detail above, since the movement of the subject is detected by using the motion vector, the device is smaller than the configuration in which the movement of the subject is detected by using the infrared camera. It is possible to reduce the cost and price.
【0038】
In addition, since the region-specific motion vector V of the region of interest R is calculated using the motion vector v of the macroblock MB1 required for image compression, the region-specific motion vector calculation unit 14 The configuration can be simplified.
【0039】
Furthermore, since the region-specific motion vector V is calculated by performing a predetermined averaging process on the motion vectors vi to vN of the macroblocks MB1 (1) to MB1 (N) belonging to the region of interest R. A highly accurate region-specific motion vector V can be obtained.
【0040】
C: Modification example In the above-described embodiment, the average motion vector v of the region of interest R is obtained, and the vector in the opposite direction of the motion vector v is defined as the region-specific motion vector V. Find the motion vector vi of the central macroblock MB1 (i), or find the most numerous motion vector v1 to vN of all macroblocks MB1 (1) to MB1 (N) belonging to the region of interest R. You may ask for it.
【0041】
In the above-described embodiment, the angle vector information S15 is obtained for each screen, but it may be obtained in a plurality of screen cycles. For example, the sum of the angle vectors for 30 screens may be output once on 30 screens as the angle vector information S15.
【0042】
In the above-described embodiment, since the motion vector V of the region of interest is calculated using the motion vector v of the macroblock MB required for image compression, there is an advantage that the configuration of the video conferencing system can be shared. , When the present invention is applied to a system or the like that does not originally have image compression, a portion for performing the same processing may be separately created.
【0043】
In the above-described embodiment, the angle of the video camera 4 is controlled, but control for moving the video camera 4 in the front-back, left-right, and up-down directions can also be added.
【0044】
[Effect of the invention]
As described above, according to the present invention, since the video camera is driven based on the motion vector of the region specified on the monitor, the speaker or the like can be used without using a special sensor such as an infrared camera. The subject can be automatically tracked. Therefore, it is extremely suitable for use in a video conferencing system or the like, and the device can be miniaturized and inexpensive.
[Simple explanation of drawings]
[Figure 1]
It is a block diagram which shows the structure of one Example of this invention.
[Figure 2]
It is a schematic diagram which shows the setting of the region of interest in the same Example.
[Fig. 3]
It is a conceptual diagram which shows the calculation process of the motion vector in the same Example.
[Explanation of symbols]
1 ... Camera movement request switch, 2 ... Camera movement request control angle calculation unit, 3 ... Camera fixed base drive unit, 4 ... Video camera, 5 ... Speaker display monitor, 6 ... AD conversion unit, 7 ... current frame buffer, 8 ... previous frame buffer, 9 ... motion vector calculation unit, 10 ... DCT / quantization unit, 11 ... entropy coding unit , 12 ... Inverse quantization / IDCT section, 13 ... Area of interest specification section, 14 ... Region-specific vector calculation section, 15 ... Control angle calculation section, 16 ... Coding section, 17. .Decoding unit, 18 ... DA conversion unit, 19 ... Other party image display monitor.
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2007088611A | Cited by | Japan | Search report |
| WO2007119355A1 | Cited by | World Intellectual Property Organization (WIPO) | Applicant |
| US8599267B2 | Cited by | United States of America | Applicant |
| US8072499B2 | Cited by | United States of America | Applicant |
| JP2009081881A | Cited by | Japan | Examiner |
| KR19990060503A | Cited by | Republic of Korea | Search report |
| US7248286B2 | Cited by | United States of America | Search report |
| KR100413268B1 | Cited by | Republic of Korea | Search report |
| KR20020095999A | Cited by | Republic of Korea | Search report |
3 priority claims, no other members on record
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 12866793 | Japan | A | |
| 5128667 | – | – | – |
| JP19930128667 | – | – | – |
Numbers
- Publication
- 6-339056
- Publication, DOCDB
- H06339056
- Publication, EPODOC
- JPH06339056
- Application
- 5128667
- Application, DOCDB
- 12866793
- Application, EPODOC
- JP19930128667
Titles3
- English
- [Title of Invention] Video camera drive device
- English
- Video camera drive
- Japanese
- 【発明の名称】ビデオカメラ駆動装置
Classification
- IPC, 3
- H04N5 232
- H04N7 15
- H04N7 18