A hand washing monitoring system
Abstract
This record has no abstract on file.
Term
0.6 yearsto projected expiry
Projected expiry 4 May 2027, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
22 claims: 8 independent, 14 dependent
- 1Zastrzeżenia patentowe 1. System kontroli mycia rąk wyposażony w kamerę (2) oraz procesor (4), przy czym procesor opracowuje uzyskane z kamery obrazy mycia rąk, znamienny tym, że za pomocą procesora analizuje się naprzemienne ruchy rąk w celu sprawdzenia, czy ręce poruszają się prawidłowo i jeżeli tak, sprawdza się czas ruchu wzorca;oraz na podstawie analizy tworzy się komunikat o jakości procesu mycia rąk.
- 2System kontroli według zastrz.1, przy czym za pomocą procesora analizuje się obrazy umieszczone w obszarze zainteresowania zawierającym złożone razem ręce.
- 3System kontroli według zastrz.1 albo 2, przy czym za pomocą procesora tworzy się komunikat niezależnie od kolejności wykonywanych ruchów.
- 4System kontroli według jednego z poprzednich zastrzeżeń, przy czym za pomocą procesora uzyskuje się informacje o właściwościach obrazów i na podstawie bazy właściwości tworzy się wektor właściwości, zawierający wektory kształtu połączonych rąk oraz przedramion, oraz uruchamia się klasyfikator z wektorami w celu określenia położenia rąk.
- 5System kontroli według zastrz.4, przy czym za pomocą procesora tworzy się na podstawie segmentacji krawędziowej co najmniej jeden wektor właściwości.
- 6System kontroli według zastrz.4 albo 5, przy czym za pomocą procesora tworzy się na podstawie pomiarów pikseli w czasie i przestrzeni co najmniej jeden wektor właściwości.
- 7System kontroli według jednego z zastrzeżeń od 1 do 6, przy czym za pomocą procesora dzieli się obraz na komórki, na podstawie których oblicza się histogramy oraz łączy się histogramy w celu utworzenia wektora właściwości.
- 8System kontroli według jednego z zastrzeżeń od 4 do 7, przy czym za pomocą procesora tworzy się wektory właściwości w oparciu o reguły wprowadzane podczas fazy treningowej za pomocą obrazów wzorcowych.
- 9System kontroli według jednego z zastrzeżeń od 4 do 8, przy czym za pomocą procesora zmniejsza się wielkość wektora właściwości.
- 10System kontroli według zastrz.9, przy czym za pomocą procesora zmniejsza się wielkość wektora właściwości stosując analizę głównych składowych oraz liniową analizę dyskryminacyjną. 1 .System kontroli według zastrz.9 albo 10, przy czym za pomocą procesora tworzy się samoorganizujący przegląd w celu przedstawienia sklastrowanych zespołów danych zredukowanego wektora właściwości.
- 1112.System kontroli według jednego z zastrzeżeń od 1 do 11, przy czym za pomocą procesora klasyfikuje się wektory właściwości przez uruchomienie klasyfikatora wieloklasowego trenowanego razem z wzorem wektora właściwości w celu oceny położenia rąk;przy czym za pomocą procesora przeprowadza się klasyfikację stosując maszynę wektorów nośnych oraz za pomocą procesora uruchamia się grupę maszyn wektorów nośnych.
- 1213.System kontroli według zastrz.12, przy czym za pomocą procesora rozpoznaje się położenie rąk na podstawie komunikatów z różnych maszyn wektorów nośnych;filtruje się również klasyfikacje położenia złączonych rąk;przeprowadza sie proces filtrowania na podstawie modeli probabilistycznych ruchów złączonych rąk w celu usunięcia wyników niezgodnych.
- 1314.System kontroli według jednego z poprzednich zastrzeżeń, przy czym za pomocą procesora rozpoznaje się czas ustawienia rąk w określonym położeniu rejestrując liczność ramek dla danego położenia oraz przyporządkowując do każdego położenia określoną minimalną wartość progową.
- 1415.System kontroli według jednego z wymienionych zastrzeżeń, przy czym podczas procesu normalizacji obrazu kalibruje się procesor ze względu na zmiany natężenie oświetlenia oraz koloru.
- 1516.System kontroli według jednego z poprzednich zastrzeżeń, przy czym za pomocą procesora przeprowadza się segmentację według koloru, struktury powierzchni oraz ruchu w celu wyeliminowania zakłóceń powodowanych przez odbicia;przy czym za pomocą procesora piksele przedstawiające biżuterię i zegarki usuwa się na podstawie analizy wielkości i kształtu obszaru,
- 1617.System kontroli według jednego z poprzednich zastrzeżeń, przy czym system jest zasilany za pomocą baterii.
- 1718.System kontroli według jednego z zastrzeżeń od 4 do 17, przy czym klasyfikator zawiera grupę słabych klasyfikatorów;i przy czym za pomocą procesora przed wyodrębnieniem właściwości rozpoznaje się skórę na podstawie koloru grup pikseli sklasyfikowanych jako skóra albo nie-skóra oraz opracowuje się dla grup liczby określające prawdopodobieństwo wyniku „skóra”.
- 1819.System kontroli według jednego z poprzednich zastrzeżeń, przy czym za pomocą procesora uruchamia się filtr kompensacji oświetlenia opracowany na zasadzie, że średni przestrzenny współczynnik odbicia dla powierzchni jest achromatyczny oraz za pomocą procesora oblicza się przepływy optyczne na podstawie wielkości i kierunku przemieszczenia pikseli i dane dotyczące ruchu przekazuje się do filtru, aby usunąć fałszywe pozytywne piksele skóry;przy czym za pomocą procesora określa się współczynniki zwiększające i zmniejszające odpowiadające zwiększeniu albo zmniejszeniu wielkości ruchu pikseli;przy czym za pomocą procesora rozpoznaje się obecność dłoni i przedramion kontrolując charakterystykę geometryczną bąbli pikseli oraz osi głównych i pomocniczych.
- 1920.System kontroli według jednego z zastrzeżeń od 4 do 19, przy czym za pomocą procesora wyodrębnia się właściwości obrazu przez tworzenie histogramów ukierunkowanych gradientów w części obrazu za pomocą gromadzenia głosów w jednej z wielu grup, gdzie jedna grupa jest przypisana do jednego kierunku gradientu;przy czym za pomocą procesora tworzy się histogram ukierunkowanych gradientów dla każdej z licznych komórek pikseli, zawierających wstępnie określone ilości tych pikseli;przy czym za pomocą procesora wszystkie histogramy łączy się tworząc jeden wektor;przy czym za pomocą procesora normalizuje sie histogramy komórek.
- 2021.System kontroli według jednego z zastrzeżeń od 4 do 20, przy czym za pomocą procesora przeprowadza się fazę treningową, w której tworzy się wektory obrazów treningowych.
- 2122. Dozownik do podawania mydła, wyposażony w system kontroli według jednego z poprzednich zastrzeżeń.
- 2223. Współpracujące z komputerem urządzenie wyposażone w kod oprogramowania do wykonywania operacji przez procesor systemu kontroli według jednego z zastrzeżeń od 1 do 21. EP 2 015 665 B1 Fig. 1 EP 2 015 665 B1 EP 2 015 665 B1 l_=Ł«L Fig. 4 EP 2 015 665 B1 Rozpoznawanie skóry Średnia wielkość ruchu Segmentowany obraz inicjalizacji Spadek Fig. 5 EP2 015 665 B1 EP 2 015 665 B1 td) Fig. 7 EP 2 015 665 B1 Operacje standardowe EP2 015 665 B1
Independent claims22
180 paragraphs in 3 sections, as filed
[0001] The present invention relates to a system for controlling hand washing.
Description of the Related Art [0002] It is well known that the spread of infection in environments such as hospitals or kitchens, resulting from improper hand washing, adversely affects health and causes major financial losses.
[0003] Employees are trained in proper hand washing in both medicine and gastronomy. Getting employees to wash their hands properly is a difficult task. Studies have shown that the fingertips and thumbs that are most often overlooked during washing are the hand parts. Especially with the help of fingers, pathogenic microorganisms spread.
[0004] The recommended hand washing time is longer than 15 seconds. Studies have shown that both long, three-minute and short ten-second hand washing can reduce the average amount of transient flora by a factor of ten and indicate that washing technique is more important than washing time.
[0005] The recommended hand washing process consists of the following phases:
1. Apply soap, moisten your hands and rub your hands.
2. Rub your back with your right hand to the level of your wrist, then repeat for the other hand.
3. Rub your left hand fingers with your right hand, repeat this step for the other hand.
4.Fingers with one hand rub the spaces between the fingers of the other hand.
5. Wash both thumbs in turn.
6. Rub the fingertips of one hand against the other hand.
7. Rinse the soap thoroughly with your hand.
8. Dry your hands without touching the tap.
[0006] WO03 / 079278f describes a system using a pattern recognition algorithm provided with a digital reference image.
This system creates binary reports informing about the effectiveness of hand washing (positive / negative result) or applying a disinfectant.
[0007] WO2005 / 093681 discloses a system for monitoring hand purity and a control method using a fluorescent indicator, an ultraviolet light source, and determining reading parameters for hands with or without an indicator.
[0008] GB2337327 describes a method based on the analysis of successive pixels for the presence of soap on the hand, a similar method is also shown in WO98 / 36258.
[0009] From US5952924, an RFID tag is known to trigger a hand washing control process. The device is activated by means of a motion detector, while the alcohol presence sensor is used to check the effectiveness of hand washing.
[0010] US6236317 describes a tracking system based on RFID technology by which it checks people entering or leaving specific rooms such as toilets or meat storage chambers. The location of the employee standing at the sink is made using active infrared detectors.
[0011] US6426701 describes a system of audiovisual instructions that guides the user through the hand washing process. To determine whether the user's hands have been in the sink for at least 20s, proximity sensors are used.
[0012] The above prior art systems do not, however, analyze the hand washing process in great detail.
[0013] The object of the invention is to improve control of the hand washing process.
Description of the idea of the invention [0014] According to the invention, a hand washing control system equipped with a camera and a processor that develops images of hand washing obtained from the camera is characterized in that the processor analyzes the alternating movements of the hands in order to check whether the hands are moving properly and if so , checks the movement time according to a pattern and, based on the analysis, creates a message about the quality of the hand washing process.
[0015] In the present invention, the processor analyzes images placed in an area of interest including hands folded together.
[0016] In another solution of the invention, the processor creates a message regardless of the order in which it moves.
[0017] In a further embodiment of the invention, information on the properties of images is obtained using a processor and a property vector is created based on the property base containing the shape vectors of joined hands and forearms, and a classifier with vectors is activated to determine the position of the hands.
[0018] In the solution of the invention, the processor based on edge segmentation forms at least one property vector.
[0019] In another embodiment of the invention, the processor creates at least one property vector based on measurements of pixels in time and space.
[0020] In another embodiment of the invention, the processor divides the image into cells on the basis of which histograms are calculated and histograms are combined to form a property vector.
[0021] The processor creates property vectors based on the rules introduced during the training phase by means of reference images.
[0022] In another embodiment of the invention, the size of the property vector is reduced by means of a processor.
[0023] In a further embodiment of the invention, the size of the property vector is reduced by means of a processor using principal component analysis.
[0024] In a preferred embodiment of the invention, the processor reduces the size of the property vector using linear discriminant analysis. [0025] In yet another embodiment of the invention, the processor creates a self-organizing overview to represent the clustered data sets of the reduced property vector.
[0026] In a further embodiment of the invention, the property vectors are classified by the processor by running the multi-class classifier trained together with the property vector pattern to assess the position of the hands.
[0027] In yet another embodiment of the invention, classification is carried out using a support vector machine.
[0028] In yet another embodiment of the invention, a group of support vector machines is started using a processor.
[0029] In a further embodiment of the invention, the processor recognizes the position of the hands by means of messages from various support vector machines.
[0030] In yet another embodiment, the position classifications of the joined hands are also filtered.
[0031] In a further embodiment of the invention, the processor performs a filtering process based on probabilistic motion models of joined hands to remove non-compliant results.
[0032] In the solution of the invention, the processor recognizes the time of placing the hands in a specific position by registering the number of frames for a given position.
In a further embodiment, the processor assigns a specific minimum threshold to each position of the hands.
[0034] In another embodiment of the invention, the processor is calibrated during image normalization due to changes in illumination and color.
[0035] In yet another embodiment of the present invention, the processor performs segmentation according to color, surface structure and motion to eliminate interference caused by reflections.
[0036] In a further embodiment, using the processor, pixels representing jewelry and watches are removed based on an analysis of the size and shape of the area.
[0037] The control system is powered by a battery.
[0038] In another embodiment of the invention, the classifier contains a group of weak classifiers.
[0039] In the solution of the invention, the processor recognizes the skin before extracting the properties based on the color of the groups of pixels classified as skin or non-skin, and numbers are developed for groups that determine the likelihood of a "skin" result.
[0040] In yet another embodiment of the invention, the processor activates a lighting compensation filter based on the principle that the average spatial reflection coefficient for a surface is achromatic.
[0041] In a further embodiment of the invention, the processor calculates optical flows based on the size and direction of pixel displacement and the motion data is passed to a filter to remove false positive skin pixels.
[0042] In another embodiment of the invention, the processor increases and decreases corresponding to an increase or decrease in the amount of pixel movement are determined using a processor, and weighting is used to detect movement or no movement.
[0043] In yet another embodiment, the processor recognizes the presence of hands and forearms by controlling the geometric characteristics of pixel bubbles as well as major and minor axes.
[0044] In another embodiment of the invention, the processor extracts image properties by creating histograms of directed gradients in a portion of the image by collecting voices for one of several groups, where one group is assigned to one gradient direction.
[0045] In another embodiment of the invention, the processor creates a histogram of directed gradients for each of the plurality of pixel cells containing predetermined amounts of these pixels.
[0046] In yet another embodiment of the invention, all histograms are combined to form one vector using a processor.
[0047] In a further embodiment of the invention, cell histograms are normalized using a processor.
In the present invention, the processor performs a training phase in which training vectors are created.
[0049] The soap dispenser is preferably equipped with the control system described above.
[0050] The solution according to the invention uses a computer-cooperating device equipped with a software code for performing operations by the processor of the control system described above.
EXACT DESCRIPTION OF THE INVENTION
Description of the drawing [0051] The invention is explained in the following part of the description on the basis of the drawing, the following figures show:
1 simplified monitoring system according to the invention;
Fig.2 block diagram of system operation;
Fig. 3 real image (left side) and aligned image (right side) at standard averaging 97;
Fig. 4 the function of increasing (left) and decreasing (right) at the threshold TID = 50;
Fig. 5 block diagram of recognition of skin presence and the movement of hands and arms;
Fig. 6 a method of selecting a controlled area;
Fig. 7 description of the directed gradient histogram (HOG), (a) real image, (b) gray image, (c) gradient image, (d) HOG cell distribution and block normalization;
Fig.8 block diagram for creating rules for criteria selection criteria and selection of these features during system operation;
Fig. 9 single result classification frame in subsequent shots; and
Fig. 10 a perspective view of a separate detector head equipped with battery power in another monitoring system.
Description of the invention [0052] The monitoring system 1 according to Fig. 1 consists of a camera 2, a light source 3, a processor 4, a data memory 5 and a display 6. The camera 2 reads the image of hand washing, and then the processor 4 processes the image by using motion tracking techniques to determine the recognition of the correct washing process.
[0053] System 1 tracks hand movements and determines in the next steps the type of movement pattern:
• Hand recognition.
• Recognizing if hands are joined.
• Defining the recognition area for joined hands.
• Inside the recognition area for joined hands, edge data and the area occupied by the hands are extracted to determine the property vectors.
• Classifying information on the movement of both hands (given in the form of property vectors) using a multi-class classifier, for example a set of support vector machine (SVM) or a cascade system of weak classifiers, to determine whether the movement of the hands corresponds to one of the standard washing movements.
• Determining whether each of the subsequent movements of the hand washing process was made in a timely manner. The order of these moves is not important. Only the set of these movements is important.
[0054] The camera 2 reads the images of washing both hands, and the processor 4 analyzes these images to determine if all the required movements have been made correctly.
[0055] The processed information is used to determine whether the hands were moving according to a specific pattern, whether the pattern time is in accordance with the specified minimum time and whether all movements have been made.
[0056] Hereinafter, the hand washing control method according to Figures 2 and 3-12 is described in detail.
Image normalization, 70 [0057] In this step, the brightness, colors and contrast of the image are modified based on the parameters measured during the subsequent stages of system startup. A typical example is determining the camera's white balance by setting a perfectly white paper in front of the camera and combining the read value with the parameters of pure white.
[0058] Algorithms for equalizing light and solid colors are used during the process. One of the uses of the invention uses a gray scale algorithm. This algorithm is based on the assumption that the average spatial reflection coefficient for a surface is achromatic. Since the light reflected from the colorless surface changes the same for all wavelengths, the spatial average of the light leaving the recognition area will be in the color of the incident lighting. The greyworld algorithm is defined as (!) = Avg avg avg where BGRc is the scale factor for each channel. BGR avg is the average value for a specific channel in a specific frame. BGRstd is 50% of the ideal canonical gray, for example BGRavg added to BGRavg = ½ BGRcanonical = 1/2256 = 128. The limitation for the Greyworld algorithm is that the standard average value of 128 cannot be matched to images with a dark background that can be overcompensated. The processor converts the standard value according to the following formulas:
(2)
G<sub>st</sub>d -
Xr.i [<sub>n</sub>. ^ (G "G" Ji) + inin (B, .O "ji)]
2xn (3)
<img file="PL2015665T3_D0001.tif" />
where m is the number of pixels in the image and n is the number of non-black pixels in the image. In this way the problem of overcompensation is solved. When calculating the average percentage of the maximum and minimum channels, an average gray value is obtained for the whole image. BGRavg is the average of non-black pixels for each channel. The scale factor Sc = Cstd / BGRavg is used for all pixels in the image.
Skin recognition, [0059] By means of a skin detection sensor, for example a Poesia filter <sup>TM</sup>, areas of the image that can be identified as the skin image are recognized. Reflective surfaces such as stainless steel can also be classified as leather with this recognition method.
[0060] Nonparametric skin modeling according to Bayesian models based on histograms is used. The color of the skin and other materials is recognized by histograms. The RGB color space is normalized based on the number of rgb and rgb stripes, with the processor counting the number of colored pixels in each Nskin (rgb) skin class as well as Nnonskin (rgb) for other materials. Finally, each strip is normalized to achieve discrete skin / non-skin conditions for the distribution of colors p (rgb / skin) / (rgb / non-skin). By marking as TS and TS all areas included in the skin and not-skin histogram, for example the number of skin and non-skin pixels in the reference set, are obtained:
(4)
<img file="PL2015665T3_D0002.tif" />
^ Onskinirgb)
T<sub>N</sub> [0061] As before:
(5)
<img file="PL2015665T3_D0003.tif" />
[0062] The Bayesian formula is therefore used to assess the skin / non-skin probability by the color of a given pixel:
(6)
P (nonskin \ rgb) = p (rgb \ skin) p (skin) p (rgb \ skiri) p (skiri) + p (rgb \ nonskin) p (nonskiń) (7)
P (nonskin \ rgb) -1 - p (skin \ rgb) [0063] The reference histograms for the skin / non-skin surface are obtained from available sources using data filtering according to the Poesia design. The project selected 323 three-dimensional RGB histograms downloaded from the Compaq database and placed the data in two files: one file contains "skin" pixels, the other file contains "not skin" pixels. The values 0.4 and 0.6 were adopted as the basic probability for the results "skin" and "not skin". After applying equation (6) for skin recognition, a probability map Xp_ (i, j) is obtained. Binary skin map Xb_ (i, j) can be made according to the selected selection threshold Th, OLThLi. The pixel is designated as the "skin" pixel if p (skin / rgb) LTh or as the "non-skin" pixel if p (skin / rgb) "Th.
Calculation by optical flow method, 72 [0064] The calculation determines the path of the pixel transition from one image to the next. The movement of each pixel is represented as the size and direction of displacement. The first object placed in the recognition area, forming the beginning of the optical flow vector, is treated as hands. Reflections and water also affect the optical flow measurement. Traffic information is calculated using dense optical flow measurement using the Lucas-Kanade method. The Coarser method, like the block fit method, does not provide enough pixel information to create reliable segmentation.
Differential method by comparison with the background, 73 [0065] A motion analysis process is also used to achieve clear segmentation. This technique involves calculating the average image containing both the skin and movement parameters. For each skin pixel that may be false positive, the processor checks the motion after applying a filter that averages the entire motion picture. If the averaged motion image is greater than the TID threshold, this means an increase in probability. If it is less than TID then the probability decreases. The increase or decrease value is calculated on the basis of the following formulas:
(8)
<img file="PL2015665T3_D0004.tif" />
(9)
<img file="PL2015665T3_D0005.tif" />
where IF and DF are the increase or decrease factors by which the processor selects how fast or how slow the changes would be if they were to move, Mi is the amount of pixel movement and and TID is the threshold value on which to decide whether to increase or decrease relative to the output image. In Fig. 4 the increasing and decreasing functions are shown for the threshold TID = 50. The ascending function reaches its maximum value quickly, while the descending function is less steep because the load either accelerates the existing movement or inhibits the movement in the direction of deceleration.
[0066] In the recognition system described, skin pixels due to the metal sink and mirror reflections are removed if the hand / arm are correctly segmented only when they are moving.
[0067] The problem of incorrect recognition of hands that have stopped moving remains unsolved. Fig. 5 schematically shows the segmentation process during the movement of the hand.
Hand and arm detection, [0068] Some misreading areas are still not recognized even after filtering in steps 71, 72 and 73. The largest objects in the image being read are analyzed to find an area where there may be hands or arms. To identify hands or arms, the shape parameters of the objects found are used, such as size, long axis position relative to the short axis, and long and short axis angle.
Image recognition of joined hands, 76 [0069] When analyzing the shape parameters of the objects formed by the hands and arms, the system decides whether the hands are connected during the rubbing movement.
High Risk Gesture Recognition, 77 [0070] If the hands are not joined, the system attempts to determine if they are placed close to the tap. This can be a high risk gesture, but only at the final stage of the washing cycle.
[0071] Controlling the movement of the hands during washing is important to ensure the required cleanliness. Equally important, however, is to prevent contamination of the hand at the end of washing, because it is using the hand rather than the elbow that the water supply is cut off. This operation can be interpreted in terms of hygiene as a beneficial effect at the beginning of washing, but unfavorable after its completion. Therefore, the detection of this activity is associated with a separate classification of risk behaviors.
Area of interest calculation, 78 [0072] Area of interest is created based on the location of the hands within the image.
[0073] After the segmentation of the palm and arm image, the processor performs a simple analysis of the elements of the binary image by removing small elements from the area of interest. At this point, the area of interest for the hands is selected. Using the vertical axis of symmetry, the upper and lower area limits are calculated, which will then be used to calculate the lateral limits. This process is illustrated in Figure 6.
[0074] The shape of the area of interest should be square in order to maintain dimensions and not to distort the position reading of the hands. Sometimes the processor increases the width or height of the area to obtain a square-shaped field. Then the dimensions of the area of interest change to form a 128 x 128 square, from which later individual properties will be obtained. Sometimes the range of the hands or their size in the image changes depending on the distance from the camera. In order to compensate for such scale changes, the processor changes the size of the area of interest according to the property recognition system, for example using the dimensions 128 x 128, 64 x 64, 32 x 32. Recognizing object properties at different scale sizes increases classification tolerance for different hand sizes placed in the image area .
Object Properties Recognition, 79 [0075] Using an algorithm such as an optical flow histogram, directional gradient histogram, or reducing the scale of a hand image, the system extracts information about the object from the image.
[0076] The selection of specific and independent properties is important for successful pattern recognition for classification purposes. While the processor uses a local directed gradient histogram (HOG) as a means of extracting the properties of a single frame, the purpose of this method is to describe the image using a set of local histograms. These histograms count the orientation of gradients in the local part of the image. First, the image gradient is calculated. This image is divided into cells, which can be defined as an area of space forming a square with a pre-selected size in pixels. For each cell, the processor calculates a histogram of gradients for each direction by collecting events. Events are weighted depending on the size of the gradient so that the histogram includes the size of the gradient at a given point.
[0077] Once all histograms for each cell have been calculated, the processor creates an image descriptor vector, combining all histograms in one vector. Due to the variability of image elements, it is necessary to normalize cell histograms. Cell histograms are locally normalized according to the values of neighbor cell histograms. Normalization is carried out in a group of cells forming a block. Then the normalization factor for the whole block is calculated and all histograms in the block are normalized according to this factor. After completing all normalizations, the histograms can be combined into a single property vector. The L2 standard scheme was used:
(10)
<img file="PL2015665T3_D0006.tif" />
[0078] Where □ is a small adjustment constant due to the need to determine sometimes empty gradients. Depending on the design of each block, the histogram from a given cell can be included in several block normalization projects. This creates redundant information that improves system performance. This method is illustrated in Fig. 7.
[0079] To calculate the vector dimension, several parameters need to be considered: ROI, cell size, block size, number of stripes, and number of overlapping blocks. The ROI dimension is 128 x 128, each of the windows is divided into 16 x 16 cells and each 2 x 2 cell group is connected to the block slidingly, so that the blocks overlap. Each cell consists of a 16-band directed gradient histogram (HOG), and each block contains a vector formed of all cells. Thanks to this, each block is represented by a 64 vector of properties, normalized to the length of the L2 unit. Each 128 x 128 area of interest is represented as:
<sup>(11) cells</sup>poems<sup>ROI =</sup>W<sup>/ cell</sup>W <sup>(12) cells</sup>column<sup>ROI =</sup>H<sup>/ cell</sup>high (13) cells-blocks = blocker / cells wide (14) cells-block-columns = block / cells high (15) overlapping-blocks = rows-cells-blocks + 1 (16) overlapping-columns = cells-columns-cells-blocks + 1 ( 17) DIM vector = overlapping blocks * overlapping block-columns * (18) cells-block lines * cells-block columns * HOG strips [0080] In this way, the dimension of the property vectors is 3.136.
[0081] With reference to Fig. 8, in a trial phase, the system uses saved gesture sequences based on observations. This information is used to sort relevant properties, train classifiers, and tune filters. In the second phase of "normal operation", the system processes the image, extracts properties, sorts properties, classifies and filters activities.
[0082] According to Fig. 3, the training phase comprises:
101, labeling images,
102, processing recorded video images,
103, property recognition,
104, reduction of dimensions,
105, creating and applying property selection rules,
106, property vector creation,
107, training classifiers and
108, setting filter rules [0083] In the training phase, a complete set of properties is created using a directed gradient histogram or flow histogram (or worm-time motion). Due to the large number of properties created, not all of them have a positive impact on classification. Of these properties, a subset of only those properties that are favorable for classifying hand movements is formed. Principal component analysis or linear discriminant analysis is used to identify this subset of properties. This process is called property selection. The subset of properties created contains about 10% of the initial set and is passed to the classifier. The classifier learns to associate movement patterns and identifies hand movements during washing. The final stage of training is the selection of filters that detect and correct classification errors by checking the sequence of hand movements by examining the sequence of hand movements based on the likelihood of different hand positions. Mathematical constructions such as Kalman filter, Baesian network or Markov decision making processes are used to implement these filters.
[0084] In the normal operating phase of the system, image processing filters are used, the area of interest (ROI) is identified, and property recognition algorithms are used. The subset of properties is selected based on the first training and developed as a property vector. This vector is passed to the classifier, which determines which hand movement is considered at a given time. The result is filtered to eliminate errors in the final classification of hand movement.
Creating the property vector, 80 [0085] An array of appropriate measurements for classification is made, usually only those measurements are used for classification purposes. Measurements obtained using property acquisition algorithms are combined to form an array called a property vector that can contain over 3,000 different measurements.
[0086] The next two steps to create a property vector are to develop a histogram for setting the edge and to develop a histogram of motion in time and space.
[0087] Flow histograms are used to create a property vector using space-time data. An area of interest is created around the arms and forearms and a high density motion vector is created. The image is divided into cells and histograms of directions of motion vector are calculated. To calculate directional histograms of a series of frames, histograms from several further frames can also be included. Histograms are normalized either globally or locally as overlapping adjacent areas. The resulting histograms are compiled to form a property vector passed to the multi-class classifier. This process has already been described above.
[0088] To reduce the dimension of the property vector, the processor analyzes the diversity of this vector using principal component analysis (PCA). This tool helps the processor 5 to identify measurement results that will be useful to obtain the correct classification. The processor retains only those properties that contribute the most to proper classification; it leaves only about 10% of the initial properties.
[0089] In another embodiment of the invention, instead of principal component analysis (PCA), tools such as linear discriminant analysis (LDA) are used, especially Fischer LDA discriminant analysis. By using smaller or larger cluster data sets, you can allow the use of simple K-nearest neighbors classifiers that can be easily inserted into the device.
Classification, 81 [0090] The property vector is passed to a multi-class classifier to develop the position of the hands at a given time. If the interference is small (good lighting, non-reflecting washbasin), you can use a simple classifier such as KNN (nearest neighbors or "superviser K-means"). If there are more problems with the overall classification, then support vector machine assembly (SVM) or a cascade of weak classifiers can be used.
[0091] The multilateral classification system is trained using a property vector to obtain reliable classification, effective computational structure, and generalization adapted to a variable video image containing the shape, size and configuration of the hand. One possible solution is to use a support vector machine assembly. Several types of classification teams can be used to implement the described strategy. One such strategy is to train individual support vector machines to classify hand positions that differ from each other. Multiple support vector machines are trained to obtain all possible combinations. The results from various support vector machines can be collated using the voting pattern in which the position of the hand with the most votes is the correct position.
[0092] More specifically, the carrier vector machine classifier is a binary classifier algorithm that looks for the optimal hyperplane as a function of decision in high-dimensional space. This is a kind of example of how to teach a machine for both classification and regression purposes. Several properties of this technique make it particularly attractive. Traditional training techniques for classifiers, such as multilayer perceptron (MLP), use empirical risk minimization and only guarantee minimal learning error. Unlike this technique, the SVM solution based on the principle of structured risk minimization minimizes the error limits and therefore gives better results.
[0093] In this way, considering one of the data sets {xk, yk} e χχ _ {-_ ι, 1}, where xk are examples of training (property vector) and yk is a class label.
The described method contains the first mapping of xk in high-dimensional space thanks to the function □. Then looking for a decision function form f_ (x) = w.
Function Φ_ (χ) + _ bi f_ (x). is optimal in the sense that it provides the maximum distance between the nearest point φ (xi) and the hyperplane. The class x label is then obtained by taking into account the sign of the function f (x).
[0094] The optimization problem can be solved for the soft margins of SVM classifiers (an example of a mistaken classification is charged linearly) as follows:
<img file="PL2015665T3_D0007.tif" />
(19) with the limit vk, ykf_ (xk)> 1 - ξk. The solution to the problem is obtained using the Lagrangian theory and can be represented as a "w" vector:
(20) m
Λ = 1 *
where is the solution to the following quadratic optimization problem:
mlm (21) maxPT (a) = "A = 1 <sup>2</sup> kj where = ° and vk, o <ak <c, where K_ (xk, xl) = (φ (xk), φ (xi) _> K_ (xk, x), is the so-called nuclear function. new kernels are proposed, beginners can find the following four primary kernels in SVM manuals:
linear Kfai, Xj) - <sup>x</sup>j Xj polynomial £ (&, xj) <sup>=</sup> (.7<sup>χ</sup>ΐ<sup>x</sup>j + rf »7 <sup>></sup> θ · radial base function (RBF) K (Xi> <sup>x</sup>j) ~ <sup>QX</sup>P <sup>- x</sup>jf<sup>2</sup>)> Z <sup>></sup> s sigmoid function
K (x<sub>b</sub> Xj) = tanh (Xj + r) where γ, rid are kernel parameters.
[0095] The LIBSVM library is used for training and testing. RBF nonlinear kernels map the sample patterns in high-dimensional space, so, unlike a linear kernel, the RBF kernel supports cases where the relationship between class labels and attributes is nonlinear. The second reason for use is the number of hyperparameters that affect the complexity of the models, because the polynomial nucleus contains more hyperparameters than RBF nuclei. In addition, the RBF kernel is easier to calculate.
[0096] To carry out the multi-class classification, the one-to-one method is used, in which k (k-1) / 2 (with k = 3 has 21 classifiers) of different binary classifiers, each of these classifiers develops data from two different classes. In practice, the one-to-one method is one of the most beneficial strategies for solving multi-class classification problems [14]. After processing the data from the classes i, j, the following selection strategy is used: if sign ((wij) T φ (x) + bij), where x is an element of the class ith, then one is added to the vote for ith. Otherwise, cfu increases by one. Then the variable x is intended for the class with the most votes. This is usually described as the "maximum winner" strategy. Where two classes have the same number of votes, the decision values (distance from the decision surface) of each classifier are taken into account.
Filtering the position of the hands 82 [0097] The position of the hands is presented as a table of transitions from the initial position with the corresponding probability assigned. For example, if the current position is B and the initial position was E, then the probability of this displacement is 80%. All current and start positions are assigned to positions determined by training and analysis. Low probability items are removed as being the result of misclassification.
[0098] The addition of additional classification filters allows the use of small and efficient classifiers. Random classification errors are captured and corrected based on a hand position probabilistic model. For example, it is unlikely that your hands will go directly from position A to position F without being in an intermediate position in the meantime. It is much more likely that position F is the result of misclassification, and the filter suggests the most likely position. Filters operate in many time intervals; the frame-to-frame transition and weighted averages over several frame intervals form the basis for creating a reliable position sequence classification for analyzing hand movements.
Time and movement analysis, 83 [0099] Positions that are fixed at a specified threshold number of frames (e.g., 3) are transferred to a battery that counts the number of frames so that each position is retained, and counts the time of each of these positions. Each position is assigned a threshold value that specifies the minimum time. If all threshold values are exceeded, the hands are considered to have been washed to an appropriately high standard.
[0100] Hand position information is used to perform a comprehensive hand hygiene assessment. Each hand position must be recorded within a set period of time. Hand movements during two-hand washing can be performed in any order. The information is related to the passage of time and transmitted to the user via a graphic display, speaker or audio device. It is possible to enter accurate motion data, for example, total washing time and critical and non-critical hand positions.
[0101] System 1 continuously graphically displays the position of the hands on the display 6. Display 6 is a low-current LCD display and informs the user which hands have been placed. This graphics can be easily adapted to any installation.
[0102] The system has to deal with problems such as, for example, lighting changes. In a typical scenario, the handwash place is placed in lighting consisting of a mixture of natural and artificial light. Each of these sources has a different color, and the proportions of natural and artificial light change depending on the time of day. The distribution of lighting is very important for reading the position of the hands, and the color of this light is an important element for the property recognition algorithm. Therefore, calibration systems for color change have been developed. They are included in the image normalization procedure.
[0103] The surfaces of many washbasins cause strong light reflections, especially the surfaces of washbasins made of stainless steel. Using only color to segment your hands may result in false messages because the color of your hands may be reflected in the surface of the sink. In segmentation based on a combination of colors, surface structure and mirror movement of reflection in typical sinks, the obtained result can be considered reliable.
[0104] If the colors and surface structure are used to separate the image of hands and forearms from the background, personal jewelry, for example watches and rings, often form an image in separated small areas.
Typically, these small areas are formed due to segmentation errors and are removed with dimensional filters. However, to determine if a small area is the result of segmentation error or jewelry recognition, the shape of the borders is also analyzed. The recorded shape of such a border is used when two shapes located on either side of the dividing line set close and parallel to the border, and in addition the color and surface structure are identical. If so, the areas are combined.
Results
Training and testing control [0105] A controlled labeling process is used to create training procedures and test data sets. First, a database is created containing images of areas of interest (ROI) and their labels. This solution allows you to use the database again. Secondly, the off-line application calculates property vectors for both training and test procedures. A classifier correctly prepared with the help of training can already be used to recognize the position of hands online. An example of the number of training sessions and tests is shown in Table 1. A separate class is assigned to each of the different hand positions during the recommended washing procedure.
Table 1. Class distribution of training and test data sets
<td>Data Sets</td><td>Position 1</td><td>Position 2</td><td>Position 3</td><td>Position 4</td><td>Position 5</td><td>Position 6</td><td>Other</td>
<td>Data Set training</td><td> 638</td><td> 780</td><td> 594</td><td> 977</td><td> 902</td><td> 1073</td><td> 587</td>
<td>Data Set test</td><td> 488</td><td> 414</td><td> 360</td><td> 456</td><td> 477</td><td> 358</td><td> 304</td>
[0106] Table 2 shows the results obtained after applying the multi-class classification using the properties of the site gradient gradient (HOG) and the SVM classifier. These are the results for individual frames, which means that they are not filtered at the processing stage. Although recognition efficiency is low for "other positions", which is justified given that it is a highly variable class, all recommended positions are classified with recognition efficiency greater than 85%. Three of these locations are classified with efficiency greater than 91%, and the best-developed class has 96.09% efficiency. A multistage process of checking the frame would significantly improve these results.
<td></td><td>Position 1</td><td>Position 2</td><td>Position 3</td><td>Position 4</td><td>Position 5</td><td>Position 6</td><td>Other</td>
<td>Effectiveness</td><td> 86,07%</td><td> 91,55%</td><td> 94,72%</td><td> 89,25%</td><td> 96,37%</td><td> 96,09%</td><td> 61,84%</td>
<td>recognition</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
[0107] Fig. 9 shows in sequence the results of a single frame classification. The designation of the recognized class is shown in the upper left corner of the image. The system correctly detects when the hands are separated or joined. The classifier even based on one frame reads the transition between subsequent positions, which are classified as different location classes. The second, fifth and sixth positions are classified correctly in both cases: the left hand rubs the right and the right hand rubs the left. [0108] Another system, 160, is shown in Fig. 10. The system 160 consists of a housing 161 in which the camera window 162 is positioned towards the sink. The processor is embedded in the housing and the system is powered by a battery. The base plate 163 is located from the outside.
[0109] By introducing an efficient cascade-based segmentation and classification system, effective but error-prone, results with high reliability are obtained. Thanks to this solution, the cost of system implementation is small, it can also work with an energy-saving computer. The system can be powered by batteries and can be integrated with a removable soap dispenser.
[0110] The system may transfer hand washing data to a storage device or this data may be transmitted via a wireless link to a central data storage database. These records may form part of the HACCP system (Hazard Analysis and Critical Control Points) used in food processing or other hygiene standards.
[0111] The scope of the invention as set out in the appended claims is not limited to the exemplary embodiment described and can be varied in design and details.
EP 2 015 665 B1
U-3310/09
Contents3
12 members in 9 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 20060350 | Ireland | A | |
| 20060350 | Ireland | A | |
| 20060905 | Ireland | A | |
| 20060905 | Ireland | A | |
| 07736107 | European Patent Office (EPO) | A | |
| 2007000051 | Ireland | W | |
| 2007000051 | Ireland | W | |
| EP20070736107 | – | – | – |
| IE20060000350 | – | – | – |
| IE20060000905 | – | – | – |
| WO2007IE00051 | – | – | – |
Members12
| Document | Office | Kind | |
|---|---|---|---|
| WO2007129289A1 | World Intellectual Property Organization (WIPO) | A1 | |
| IE20070336A1 | Ireland | A1 | |
| EP2015665A1 | European Patent Office (EPO) | A1 | |
| US2009087028A1 | United States of America | A1 | |
| EP2015665B1 | European Patent Office (EPO) | B1 | |
| AT439787T | Austria | T | |
| ATE439787T1 | Austria | T1 | |
| DE602007002068D1 | Germany | D1 | |
| DK2015665T3 | Denmark | T3 | |
| ES2330489T3 | Spain | T3 | |
| PL2015665T3This record | Poland | T3 | |
| US8090155B2 | United States of America | B2 |
Numbers
- Publication, DOCDB
- 2015665
- Publication, EPODOC
- PL2015665T
- Application
- 736107
- Application, DOCDB
- 07736107
- Application, EPODOC
- PL20070736107T
Titles2
- English
- A HAND WASHING MONITORING SYSTEM
- Polish
- System do kontrolowania mycia rąk
Classification
- CPC, 2
- G08B21/245
- G06V40/28
- IPC, 4
- A47K5 00
- G06K9 00
- G08B21 24
- H04N23 40