Concurrent learning and performance information processing system
Abstract
(57) [Summary] At the beginning of each time trial, a vector of measurements and a vector of measurement ambiguities are given to the system (10), and learning weights are given to or generated from the system (10). System (10) then performs the following operations during each time trial. Convert measurements to feature values, convert measurement ambiguity values to feature viability values, use each viability value to determine the vanishing value state of each feature value, use non-disappearing feature values for parameters Update learning, pass on each vanishing feature value from non-disappearing feature values and / or pre-learning, convert the passed-through feature value to an output pass-through measurement, apply a variety of feature values and feature function monitoring and interpretation statistics .. In the parallel embodiment of system (10), all such operations are performed in parallel by the coordinated use of the parallel feature processor (31) and the joint access memory (23), which connects the feature processors in pairs, including join weights. Will be executed.
Term
Term ended
Projected expiry passed 1 November 2015, 10.9 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
1 claim: 1 independent, 0 dependent
- 1【特許請求の範囲】 1. タイムトライアル中に受信した入力値から出力値を計算する並列処理シス テムであって、該システムは、 各々がタイムトライアル中に並列に入力ベクトルからの個別の入力値を受信す るように作動する複数の処理装置と、 前記各処理装置を前記システムの1台おきの処理装置に接続し前記プロセッサ 間で重み付けされた値を転送するように作動する複数の相互接続された導体とを 具備し、 前記各処理装置前記タイムトライアル中に前記重み付けされた値に基づいて予 期された出力値を与えるように作動しかつ前記タイムトライアル中に前記入力値 に基づいて前記重み付けされた値を更新するように作動する並列処理システム。 2. 請求項1記載の装置であって、さらに前記相互接続された導体に沿って配 置された複数のスイッチングジャンクションを具備し、前記スイッチングジャン クションは前記各プロセッサを前記システムの1台おきのプロセッサと一意的に 対とする装置。 3. 請求項2記載の装置であって、前記各スイッチングジャンクションはタイ ムトライアル中に前記各プロセッサを前記プロセッサの他の1台だけと選択的に 接続し、前記プロセッサの多数の対を形成して、時間間隔中に前記重み付けされ た値を伝達する装置。 4. 請求項3記載の装置であって、さらに複数のメモリ素子を具備し、前記各 メモリ素子は別々のスイッチングジャンクションに個別に接続されておりかつ重 み付けされたメモリ値を含む装置。 5. 請求項3記載の装置であって、前記スイッチングジャンクションは多数の タ時間間隔中に前記プロセッサの多数の対の中の異なる対を選択的に接続する装 置。 6. 請求項5記載の装置であって、前記スイッチングジャンクションは最小数 のステップで前記プロセッサの多数の対の中の前記異なる対の起こり得る全ての 組合せを選択的に接続する装置。 7. 請求項6記載の装置であって、さらに複数のメモリ素子を具備し、前記各 メモリ素子は別々のスイッチングジャンクションに個別に接続されておりかつ重 み付けされたメモリ値を含む装置。 8. 請求項3記載の装置であって、前記導体は第1の導体層および第2の導体 層を具備し、前記第1および第2の導体層は前記スイッチングジャンクションに おいて接続するように作動する装置。 9. 請求項8記載の装置であって、前記導体は第1の導体層および第2の導体 層を具備し、前記第1および第2の導体層は前記スイッチングジャンクションに おいて接続するように作動する装置。 10.請求項9記載の装置であって、前記スイッチングジャンクションは半導体 層内に配置され、前記第1および第2の導体層間に配置されている装置。 11.請求項10記載の装置であって、さらに複数のメモリ素子を具備し、前記 各メモリ素子は別々のスイッチングジャンクションに個別に接続されておりかつ 重み付けされたメモリ値を含む装置。 12.請求項11記載の装置であって、前記各スイッチングジャンクションは前 記各プロセッサを前記プロセッサの他の1台だけと選択的に接続し、前記プロセ ッサの多数の対を形成して、時間間隔中に重み付けされた値を伝達する装置。 13.請求項1記載の装置であって、さらに外部測定ベクトルを受信して前記入 力ベクトルへ変換するトランスジューサ入力プロセッサを具備し、前記トランス ジューサ入力プロセッサは時間順とされた予め選定された数の前記入力ベクトル を格納するように作動する装置。 14.タイムトライアル中に受信する入力値から出力値を計算する処理システム であって、該システムは、 入力ベクトルから、逐次、入力データ値を受信するように作動する処理装置と 、 データ列として逐次格納された結合重みマトリクスの要素を含む前記処理装置 に接続された記憶装置とを具備し、 前記処理装置は、前記タイムトライアル中に、前記結合重みマトリクスの前記 要素に基づいて予期された出力値を与えかつ、前記タイムトライアル中に、前記 入力データ値に基づいて前記結合重み値を更新するように作動する処理装置。 15.請求項14記載の装置であって、前記プロセッサは前記結合重みマトリク スの各要素へ逐次アクセスするように作動する装置。 16.請求項15記載の装置であって、前記プロセッサは前記列内で現在アクセ スされている結合重みに出会う時にそれに基づいて全ての転嫁オペレーションを 実行するように作動する装置。 17.請求項16記載の装置であって、前記プロセッサは前記データ列内で現在 アクセスされている結合重みに出会う時にそれに基づいて全ての更新オペレーシ ョンを実行するように作動する装置。 18.多数のタイムトライアル中に与えられる入力データベクトルm[IN]内 に含まれる多数の入力値[解析入力データ]間の統計的関係を識別するコンピュ ータシステムであって、該コンピュータシステムは、 タイムトライアル中に入力ベクトルを受信するように作動する処理装置と、 受信した事前m[IN]ベクトルに基づいて現在の入力m[IN]要素間の関 係を表す結合重み要素を含む記憶装置とを具備し、 前記処理装置は前記受信した入力データベクトルの非消失値に基づいて前記結 合重み要素を更新しかつコンポーネント学習重みl[C](f)に基づいて前記 結合重み要素を更新するように作動し、前記l[C](f)は受信した各m[I N](f)データベクトルの異なる学習重みであって前記入力データベクトル要 素m[IN](f)が受信した事前測定ベクトル要素m[IN](f)に対して 行う前記結合重み要素の調整の量を決定するコンピュータシステム。 19.タイムトライアル中に受信する入力データベクトルから出力値を計算する 並列処理システムであって、該システムは、 タイムトライアル中に受信する前記入力データベクトルの別々の値を処理する 並列接続された複数の処理装置と、 前記各プロセッサに接続され要素のマトリクスの逆元である結合重みマトリク スを含む憶装置とを具備し、 前記各処理装置は、前記トライアル中に、前記結合重み要素を使用して数学回 帰解析を実施して消失入力データ値の予期された出力値を与えかつ、前記トライ アル中に、現在のトライアルからの入力データベクトル要素の関係を反映するよ うに前記各結合重み値を直接更新するように作動する並列処理装置。 20.複数の処理装置の各々間に通信経路を提供する装置であって、該装置は、 第1のプロセッサ、第2のプロセッサおよび第3のプロセッサと、 前記第1のプロセッサに接続された第1の導体と、 前記第2のプロセッサに接続された第2の導体と、 前記第3のプロセッサに接続された第3の導体と、 各々が前記第1の導体に沿った異なる点に接続された第1および第2のスイッ チングジャンクションとを具備し、 前記第2の導体経路は前記第2のプロセッサから前記第1のジャンクションへ 延在していて、前記第1のジャンクションは前記第2の導体経路および前記第1 の導体経路を介して前記第1のプロセッサを前記第2のプロセッサへ接続するよ うに作動し、 前記第3の導体経路は前記第3のプロセッサから前記第2のスイッチングジャ ンクションへ延在していて、前記第2のスイッチングジャンクションは前記第3 の導体経路および前記第1の導体経路を介して前記第1のプロセッサを前記第3 のプロセッサへ接続するように作動し、 第3のジャンクションは前記第3の導体経路および前記第2の導体経路に接続 されていて、前記第3の導体経路および前記第2の導体経路を介して前記第3の プロセッサを前記第2のプロセッサへ接続するように作動する装置。 21.請求項20記載の装置であって、さらに複数のメモリ素子を具備し、前記 各スイッチングジャンクションに1個の前記メモリ素子が配置され、前記メモリ 素子は前記第1、第2および第3のスイッチングジャンクションに接続された前 記プロセッサによりアクセスすることができる装置。 22.請求項21記載の装置であって、前記メモリ素子は重み値を含む装置。 23.請求項22記載の装置であって、さらに前記スイッチングジャンクション のスイッチングをコントロールするように作動するコントロールプロセッサユニ ットを具備する装置。 24.請求項23記載の装置であって、前記コントロールユニットは前記スイッ チングジャンクションへ第1の信号を与えて各処理装置を前記メモリ素子の選定 された素子へ選択的に接続するように作動する装置。 25.請求項24記載の装置であって、前記コントロールユニットは前記各スイ ッチングジャンクションへ第2のコントロール信号を与えて前記各処理装置を前 記スイッチングジャンクションに接続された他方のプロセッサユニットに接続す るように作動する装置。 26.請求項25記載の装置であって、前記コントロールユニットは前記スイッ チングジャンクションへ第3のコントロール信号を与えて前記スイッチングジャ ンクションに接続された前記プロセッサの1台を前記メモリ素子の選定された素 子へ選択的に接続するように作動する装置。 27.多数のタイムトライアル中に与えられる入力デー入力ベクトルm[IN] を解析することにより多数の変数間の統計的関係を識別してm[IN](f)の 消失データ値の予期された出力、m[OUT](f)、を反映する出力ベクトル を与えるコンピュータシステムであって、該システムは、 タイムトライアル中に前記処理装置において入力データベクトルを受信するス テップと、 前記タイムトライアル中に、共変マトリクスの逆元ω[IN]を形成する結合 重み要素に基づいて前記データベクトルの消失入力値に対する出力値を転嫁する ステップと、 前記現在のタイムトライアル中に、受信した前記入力測定ベクトルの非消失値 に基づいて前記結合重み要素を更新し、更新した逆元共変マトリクスω[OUT ]を形成するステップとを含むコンピュータシステム。 28.請求項27記載の方法であって、さらにコンポーネント学習重みl[C] (f)に基づいて前記結合重み要素を更新するステップを含み、前記l[C]( f)は受信した各測定ベクトルの異なる学習重みであって、前記更新ステップ中 に、前記入力データベクトルが受信した事前入力データベクトルに対して行う調 整の量を決定する方法。 29.請求項27記載の方法であって、さらに受信した前記入力データベクトル を含む全ての事前データベクトルの事前平均ベクトルに基づいて前記結合重みマ トリクスを更新するステップを含む方法。 30.請求項28記載の方法であって、さらに受信した全ての事前測定値ベクト ルベクトルの事前平均ベクトルに基づいて前記結合重みマトリクスを更新するス テップを含む方法。 31.請求項28記載の方法であって、該方法はさらに、 前記更新ステップ中に前記入力データベクトルが受信した事前データベクトル に対して行う調整の量を決定するのに使用するグローバル学習重み1を受信する ステップと、 各入力データベクトルの事前該重みおよび消失値履歴のインジケータである学 習履歴パラメータベクトルλ[IN]を受信するステップと、 入力データベクトルの消失の程度を示すバイアビリティベクトルを受信するス テップと、 方程式1[C](f)=lν(f)λ[IN]に従って前記各コンポーネント 学習重み1[C](f)を計算するステップとを含む方法。 32.請求項31記載の方法であって、該方法はさらに、コンポーネント学習重 みl[C](f)および事前学習履歴パラメータλ[IN]に基づいて前記学習 履歴パラメータベクトル要素を更新するステップを含み、 である方法。 33.請求項31記載の方法であって、該方法はさらに、 受信した全ての事前測定値ベクトルの事前平均ベクトルμ[IN]に基づいて 前記結合重みマトリクスを更新するステップを含み、前記事前平均ベクトルμ[ IN]は前の測定トライアルからのμ[OUT]に等しくかつm[OUT]の各 要素は次式に従って計算され、 ここに、μ[IN](f)は最初の測定トライアルに対しては1に等しい方法。 34.請求項33記載の方法であって、該方法はさらに、中間転嫁ベクトルe[ IN]を利用して前記結合重み要素を更新するステップを含む方法。 35.請求項34記載の方法であって、e[IN]の各要素は次式に従って計算 される方法。 36.請求項35記載の方法であって、該方法はさらに、下記の計算プロセスに 従って前記結合重みマトリクスを更新するステップを含み、 ここに、 かつ d=e[IN]ω[IN]e[IN] T =χe[IN] T であり、前記ω[IN]は前の測定トライアルからのω[OUT]に等しくかつ 最初のタイムトライアルの前の恒等マトリクスに等しい方法。 37.請求項35記載の方法であって、該方法はさらに、消失値m[IN](f )を転嫁する前記ステップを含み、 m[OUT](f)=μ[IN](f)+e[IN](f)(2-ν(f)) +χ(f)(ν(f)-1)/ω(f,f) である方法。 38.請求項35記載の方法であって、該方法はさらに、受信した全ての事前デ ータベクトルの事前平均ベクトルμ[IN]に基づいて分散ベクトルν[D,O UT]を計算するステップを含み、前記事前平均μ[IN]は前の測定トライア ルからのμ[OUT]に等しくν[D,OUT]の要素は次式に従って計算され 、 ここに、ν[D,IN]は前のトライアルからのν[D,OUT]に等しく最初 のトライアルについては1に等しい方法。 39.請求項38記載の方法であって、該方法はさらに、非消失学習平均値から 標準二乗偏差値d[1](f)を計算するステップを含み、 d[1](f)=(m[IN]-μ[OUT](f)) 2 /ν[D,OUT] (f) である方法。 計算するステップを含み、ここに、 である方法。 41.請求項40記載の方法であって、該方法はさらに、受信した全ての事前測 定値ベクトルの事前平均ベクトルμ[IN](f)に基づいて分散ベクトルν[ D,OUT](f)を計算するステップを含み、前記事前平均μ[IN](f) は前の測定トライアルからのμ[OUT](f)に等しく, であり、ν[D,IN](f)は前のトライアルからのν[D,OUT](f) に等しく最初のトライアルについては1に等しい方法。 42.請求項40記載の方法であって、該方法はさらに、回帰値d[2](f) かんの標準二乗偏差値d[2](f)を計算するステップを含み、ここに、 である方法。 43.請求項36記載の方法であって、さらにχを計算するための1組のプロセ ッサ対をアクセスするステップを含み、 前記プロセッサ対はスイッチングジャンクションに接続されており、前記各ス イッチングジャンクションは1対のプロセッサにしか接続されずタイムトライア ルi中に前記各プロセッサfを前記システムの1つおきのプロセッサgと一意的 に対とするように作動し、前記各スイッチングジャンクションは前記結合重みマ トリクスω[IN]の1つの要素ω[IN](f,g)に接続されており、該方 法は、 (a)時間間隔中に多数のプロセッサセットにアクセスするステップであって 、前記各プロセッサは前記時間間隔中に他の1台のプロセッサとしか対とされな い前記ステップと、 (b)前記スイッチングジャンクションに位置する結合重み要素ω[IN]( f,g)を各プロセッサにより検索するステップと、 (c)前記スイッチングジャンクションに接続されたプロセッサgにはe[I N],(f)をプロセッサ(f)にはe[IN](g)を転送するステップと、 (d)各プロセッサf内でx(f)=x+e[IN](g)ω[IN](f, g)を計算しプロセッサg内でx(g);+e[IN](f)ω[IN](f, g)を計算するステップと、 (e)全てのプロセッサが前記システムの1つおきのプロセッサと対とされる までステップ(a)から(d)を繰り返すステップとを含む方法。 44.請求項36記載の方法であって、さらにω[IN]の要素を更新するため の1組のプロセッサ対をアクセスするステップを含み、 前記プロセッサ対はスイッチングジャンクションに接続されており、前記各ス イッチングジャンクションは1対のプロセッサにしか接続されず選定された時間 間隔中に前記各プロセッサfを前記システムの1つおきのプロセッサgと一意的 に対とするように作動し、前記各スイッチングジャンクションは前記結合重みマ トリクスω[IN]の1つの要素ω[IN](f,g)に接続されており、該方 法は、 (a)タイムトライアル中に多数のプロセッサセットにアクセスするステップ であって、前記各プロセッサは前記時間間隔中に他の1台のプロセッサとしか対 とされない前記ステップと、 (b)前記スイッチングジャンクションに位置する結合重み要素ω[IN]を 各プロセッサにより検索するステップと、 (c)前記スイッチングジャンクションに接続されたプロセッサgにはe[I N],(f)をプロセッサ(f)にはe[IN](g)を転送するステップと、 (d)各プロセッサf内でx(f)=x+e[IN](g)ω[IN](f, g)を計算しプロセッサg内でx(g);+e[IN](f)ω[IN](f, g)を計算するステップと、 (e)全てのプロセッサが前記システムの1つおきのプロセッサと対とされる までステップ(a)から(d)を繰り返すステップとを含む方法。 45.請求項36記載の方法であって、さらに単一のプロセッサコンピュータシ ステムでχを計算するステップを含み、 前記単一のプロセッサはω[IN]の要素をその連続列としてデータ構造に格 納し、該方法は、 (a)データ列の結合重み要素ω[IN]にアクセスするステップと、 (b)結合重みマトリクスからのω[IN]要素のロー値に対応するxを更新 するステップと、 (c)列内の次のω[IN]へアクセスするステップと、 (c)現在のω[IN]要素が共変マトリクスの主対角線上にはなくかつ結合 重みマトリクスからのω[IN]要素のカラム値に対応するxを更新する最後の ω[IN]ではない場合に、 (d)ステップ(a)から(c)を繰り返すステップを含む方法。 46.請求項36記載の方法であって、さらに単一のプロセッサコンピュータシ ステムでω[IN]を更新するステップを含み、 前記単一のプロセッサはω[IN]の要素をその連続列としてデータ構造に格 納し、該方法は、 (a)データ列の第1の結合重み要素ω[IN]にアクセスするステップと、 (b)ω[IN]を更新するステップと、 (c)次のω[IN]へアクセスするステップと、 (c)現在のω[IN]要素が共変マトリクスの主対角線上にありかつ最後の ω[IN]ではない場合にステップ(b)へ進むステップを含む方法。 47.入力データを解析するコンピュータシステムにおいて、多数のタイムトラ イアル中に与えられる入力データベクトルm[IN]を解析することにより多数 の変数間の統計的関係を識別する方法であって、該方法は、 前記処理装置においてタイムトライアル中に入力データ装置から入力データベ クトルを受信するステップと、 共変マトリクスの逆元である結合重みマトリクスω[IN]の結合重み要素を 検索するステップと、 受信した前記入力データベクトルの非消失値に基づいて前記結合重みマトリク ス要素を更新しかつコンポーネント学習重みl[C](f)に基づいて前記結合 重み要素を更新するステップであって、前記l[C](f)は受信した各測定ベ クトルに対する異なる学習重みであって前記測定ベクトルが受信した事前測定ベ クトルに対して行う調整の量を決定し、更新した逆元共変マトリクスω[OUT ]を形成するステップとを含むコンピュータシステム。 48.請求項47記載の方法であって、該方法はさらに、受信した前記入力デー タベクトルを含む全ての事前測定値ベクトルの事前平均ベクトルに基づいて前記 結合重みマトリクスを更新するステップを含む方法。 49.請求項48記載の方法であって、該方法はさらに、受信した全ての事前デ ータベクトルの事前平均ベクトルに基づいて前記結合重みマトリクスを更新する ステップを含む方法。 50.請求項47記載の方法であって、該方法はさらに、 前記更新ステップ中に前記データベクトルが受信した事前データベクトルに対 して行う調整の量を決定するのに使用するグローバル学習重み1を受信するステ ップと、 各入力データベクトルの事前学習重みのインジケータである学習履歴パラメー タλ[OUT](F)を受信するステップと、 入力測定ベクトルの消失程度を示すバイアビリティベクトルν(f)を受信す るステップと、 式に従って前記コンポーネン 51.請求項50記載の方法であって、該方法はさらに、コンポーネント学習重 学習履歴パラメータベクトルを更新するステップを含み、前記事前学習履歴パラ メータλ[IN]は前のトライアルからのλ[OUT]に等しく、 である方法。 52.請求項50記載の方法であって、該方法はさらに、 受信した全ての事前測定値ベクトルの事前平均ベクトルμ[IN]に基づいて 前記結合重みマトリクスを更新するステップを含み、前記事前平均μ[IN]は 前の測定トライアルからのμ[OUT]に等しくμ[OUT]の各要素は次式に 従って計算され、 ここに、μ[IN]は最初のタイムトライアルの前はゼロに等しい方法。 53.請求項52記載の方法であって、該方法はさらに、中間転嫁ベクトルe[ IN]を利用して前記結合重み要素を更新するステップを含む方法。 54.請求項53記載の方法であって、e[IN]の各要素は次式に従って計算 される方法。 55.請求項54記載の方法であって、該方法はさらに、下記の計算プロセスに 従って前記結合重みマトリクスを更新するステップを含み、 ここに、 かつ d=e[IN]ω[IN]e[IN] T =χe[IN] T であり、前記ω[IN]は前の測定トライアルからのω[OUT]に等しくかつ 最初のタイムトライアルの前の恒等マトリクスに等しい方法。 56.請求項55記載の方法であって、さらにχを計算する1組のプロセッサ対 をアクセスするステップを含み、 前記プロセッサ対はスイッチングジャンクションに接続されており、前記各ス イッチングジャンクションは1対のプロセッサにしか接続されずタイムトライア ルi中に前記各プロセッサfを前記システムの1つおきのプロセッサgと一意的 に対とするように作動し、前記各スイッチングジャンクションは前記結合重みマ トリクスω[IN]の要素ω[IN](f,g)に接続されており、該方法は、 (a)多数組みのプロセッサを時間間隔中にアクセスするステップであって、 前記各プロセッサは前記時間間隔中に他の1台のプロセッサとしか対とされない 、前記アクセスステップと、 (b)前記スイッチングジャンクションに位置する結合重み要素ω[IN]( f,g)を各プロセッサにより検索するステップと、 (c)前記スイッチングジャンクションに接続されたe[IN]を転送するス テップと、 (d)各プロセッサf内でx(f)=x+e[IN](g)ω[IN](f, g)を計算しプロセッサg内でx(g);+e[IN](f)ω[IN](f, g)を計算するステップと、 (e)全てのプロセッサが前記システムの1台おきのプロセッサと対とされる までステップ(a)から(d)を繰り返すステップを含む方法。 57.請求項55記載の方法であって、さらにω[IN]の要素を更新する1組 のプロセッサ対をアクセスするステップを含み、 前記プロセッサ対はスイッチングジャンクションに接続され、前記各スイッチ ングジャンクションは1対のプロセッサとしか接続されず選定された時間間隔中 に前記各プロセッサ(f)を前記システムの1台おきのプロセッサ(g)と一意 的に対とするように作動し、前記各スイッチングジャンクションは前記結合重み マトリクスω[IN]の要素ω[IN](f,g)に接続されており、該方法は 、 (a)多数組みのプロセッサを時間間隔中にアクセスするステップであって、 前記各プロセッサは前記時間間隔中に他の1台のプロセッサとしか対とされない 、前記アクセスステップと、 (b)前記スイッチングジャンクションに位置する結合重み要素ω[IN] (f,g)を前記プロセッサの1台により検索するステップと、 (c)前記結合重み要素ω[IN](f,g)を更新するステップと、 (d)前記結合重み要素ω[IN](f,g)を転送するステップと、 (e)全てのプロセッサが前記システムの1台おきのプロセッサと対とされる までステップ(a)から(d)を繰り返すステップを含む方法。 58.請求項55記載の方法であってさらに単一のプロセッサコンピュータシス テムでχを計算するステップを含み、 前記単一のプロセッサはω[IN]の要素をω[IN]要素の連続列としてデ ータ構造に格納し、該方法は、 (a)前記データ列のω[IN]の結合重み要素にアクセスするステップと、 (b)結合重みマトリクスからのω[IN]要素のロー値に対応するx(f) を更新するステップと、 (c)列内の次のω[IN]へアクセスするステップと、 (d)現在のω[IN]要素が共変マトリクスの主対角線上にはなくかつ結合 重みマトリクスからのω[IN]要素のカラム値に対応するx(g)を更新する 最後のω[IN]ではない場合に、 (e)ステップ(a)から(c)を繰り返すステップを含む方法。 59.請求項55記載の方法であってさらに単一のプロセッサコンピュータシス テムでω[IN]を更新するステップを含み、 前記単一のプロセッサはω[IN]の要素をその連続列としてデータ構造に格 納し、該方法は、 (a)データ列の最初の結合重み要素ω[IN]にアクセスするステップと、 (b)ω[IN]を更新するステップと、 (c)次のω[IN]要素へアクセスするステップと、 (c)現在のω[IN]要素が共変マトリクスの主対角線上にありかつ最後の ω[IN]ではない場合に、ステップ(b)へ進むステップを含む方法。 60.入力測定ベクトルから出力値を計算して前記システムの入力および出力測 定値を評価する情報処理システムであって、該システムは、 処理装置の第1のサブシステムであって初期入力測定値を受信し、入力変換関 数に従って、前記入力測定値を前記第1のサブシステムが使用する入力特徴値へ 変換して非消失入力特徴値および/もしくは入力学習回帰パラメータから出力特 徴値を転嫁し、前記第1のサブシステムは前記出力特徴値を最終出力測定値へ変 換するように作動し、学習および出力性能の基礎となる結合重みの記憶装置を含 む前記第1のサブシステムと、 前記第1のシステムに接続されてそこから出力データを受信してディスプレイ し評価する処理装置の第2のサブシステムとを具備する情報処理システム。 61.請求項60記載の情報処理システムであって、前記第2のサブシステムは さらに前記第1のサブシステムの学習機能をコントロールするように作動する第 1のコントローラを具備する情報処理システム。 62.請求項60記載の情報処理システムであって、前記第2のサブシステムは さらに入力変換機能および出力変換機能をコントロールするように作動する第2 のコントローラを具備する情報処理システム。 63.請求項60記載の情報処理システムであって、前記第2のサブシステムは 前記メモリから前記結合重み要素を受信するように作動する情報処理システム。 64.請求項63記載の情報処理システムであって、前記第2のサブシステムは 前記結合重みを受信して前記調整された重みを前記ディスプレイおよび評価のた めに転送するように作動する情報処理システム。 65.請求項60記載の情報処理システムであって、前記第2のサブシステムは 入力値の異常な偏差が生じる場合に前記第1のサブシステムの学習機能をディセ ーブルするように作動する情報処理システム。 66.請求項65記載の情報処理システムであって、前記第2のサブシステムは 現在の測定値と前のトライアルから格納された対応する測定値間の1次差を計算 して突然変化を識別する情報処理システム。
13 paragraphs, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
Simultaneous learning and performance information processing system Field of invention In general, the present invention particularly relates to a series of parallel processing niuro computer systems. For real-time parallel processing where learning and performance occur during measurement trials Related. background Traditional statistical software and traditional neural network software are learning I / O relationships are identified during the operation, and the learned I / O relationships are during the performance phase. Apply to. For example, during the learning phase, the neural network has a well-known target from a well-known input value. Adjust the join weights until the output value is generated. God during the performance phase The network is not from known input values using the coupling weights identified during the learning phase. Extract the output value of knowledge. Traditional neural networks consist of simple interconnected processing elements. Each processing element The basic operation of the child is to convert the input signal into a useful output signal. Each interconnection The signal is transmitted from one element to another, and the relative effect on the output signal is specific. Depends on the weight of the interconnect. Traditional neural networks have well-known input and output values Can be given to the net for learning, which changes the weight of the interconnect. Various conventional neural network learning methods and models are open for mass parallel processing It has been issued. The backpropagation method is the most widely used learning method. The Multilayer Perceptron is the most widely used model. Multi-layer Perceptro The device has two or more processing element layers, most commonly an input layer and one concealing layer. It is a power layer. The concealment layer allows conventional neural networks to identify non-linear I / O relationships Includes processing elements to be processed. The learning and performance operation of the conventional neural network is performed by the neural network processing element. Since it can be operated in parallel, it can be carried out quickly during each stage. Conventional neural circuit The accuracy of the network is a user-specified date, including the number of layers and the number of processing elements in each layer. Depends on the predictability and network structure of the data. Traditional neural network learning occurs when a set of learning records is attached to the network. Each such record contains a fixed input value and an output value. The net is each record First with the join weight and other parameters learned up to that point Both update network learning by calculating network output as a function of record input .. Next, the weight is adjusted according to the approximation between the calculated output value and the learning record output value. Is done. For example, assume that the learned output value is 1.0 and the network calculation value is 0.4. Mistake The difference is 0.6 (1.0-0.4 = 0.6), which minimizes the error Used to find the weight adjustment required for. Learning records like this are used Learning is performed by adjusting the weights in the same way until then, and then all errors are sufficient. The process is repeated until reduced to. Traditional neural network learning and performance phases differ in two fundamental ways. It has become. The weight value changes during training to reduce the error between training and computational output. However, the weight value is constant during the performance phase. In addition, the output value is learned Well known during the phase, but only predictable during the performance phase .. Expected output values are learned during the performance phase input values and learning phase It is a function of the learned join weight value. Input / output relationship identification by conventional statistical analysis and neural network analysis is for some applications Satisfactory, but limited in usefulness for different applications of both methods Will be done. Extensive learning and experience, and time to perform effective manual data analysis It requires time-consuming efforts. Learning and effort required by traditional neural network analysis Less powerful, but the results obtained are less reliable and more interpretable than manual results Have difficulty. Defects in conventional statistical methods and conventional neural network methods are carried out by each method. It results from different learning and performance phases. Two different Considerable learning time before starting performance due to the need for phases Spent. The manual statistical method is an analysis of trained experts. However, since it takes a considerable amount of time, learning delays occur, and the neural network method requires enormous learning. A learning delay occurs because it requires many learning paths through the learning record. But Therefore, conventional statistical analysis is based on (a) learning occurrence time and learning model usage time. A delay is allowed between (b) the start of the time learning analysis and the start of the performance operation. It is limited to the situation where the input / output relationship with and is stable. Therefore, the information processing system of the prior art can be studied within any time trial. There is a need for rapid learning and / or performance. Abstract of the invention In general, the invention receives the measured input value of a variable during a time trial. Gradually (learn) the relationships between variables by improving the relationships learned for each trial It provides a data analysis system. In addition, any input value disappears If so, the present invention relates to previously learned relationships between the analyzed variables during a time trial. Gives the expected (imputed) output value of the lost value based on. In particular, the present invention executes a mathematical regression analysis feature value, which is a predetermined function of the input value, and transforms it. Give the bride price. Regression analysis uses a matrix of join weights to convert each feature value to another feature This is done by predicting as a weighted sum of the values. Each coupling weight element Reflect new join weight information from trial input measurements during trial Will be updated. Also, the impact of the input measurement vector on the previously received vector The component learning weights for determining the amount of data are also used during each trial. In the examples With respect to this, the present invention can process input values in parallel or sequentially. Various input values can be given in vector format. Input feature value vector Each value is calculated individually for the previously learned parameters. In the parallel embodiment, multiple A number of processors process the input value, and each processor has a specific input value from the vector. Is for reception only. That is, the system has 16 feature values (ie, 16 lengths). Processes each input feature value when set to receive (corresponding to the vector of) 16 processing devices are used for this. In the example of sequential processing, each input feature value One processor is used to process sequentially. In an embodiment of parallel processing of the present invention, each processing apparatus inputs during a time trial. Acts to receive individual input values from the vector. Each place by multiple conductors The physical equipment is connected to every other processing equipment in the system. Conductors are the process of the present invention The weighted value is transferred between the processing devices according to the above. Each processing device is described in Thailand. Gives an output value to pass on based on the weighted value during the trial. Also , During the same time trial, each processor is weighted based on the input value received Update the join weights to calculate the value. For data processing due to the limited number of outputs that a particular processor can drive When many processing devices are interconnected in parallel, one processing device interconnects them. The number of processing devices that can be driven is substantially limited. However, according to the present invention To alleviate problems related to one processor that has been contacted by many processors Multiple switching junctions along conductors interconnected with are provided. Switching junctions make each processor every other processor in the system Operates to be uniquely paired with. Further switching junk according to the present invention A memory element connected to the device is provided. Each memory element is an independent switch It is individually connected to the junction and contains one join weight value. Preferably The coupling weight memory element located at the switching junction is the total output value. It is a coupling weight element of the matrix used for calculation. Switching junctions are every other processor at a time A pair of processors that selectively connect and communicate with the weight value during the time interval Can be actuated to form. Preferably, a switching junk On a large number of time intervals, the various sets of the large number of processor pairs are continuously displayed. Can be connected to. Also, preferably there are many switching junctions. Touch various pairs of processors in every combination that can be the minimum number of steps It can be operated to continue. Control unit is heavy between processors Switch to a switching junction to control the transfer of the found value It operates to give a signal. Conductors through which processor communication takes place are preferably It is provided in the first conductor layer and the second conductor layer, and the first and second conductor layers are switched. It can be operated to connect at a ching junction. The present invention can also be implemented in a sequential manner, where conventional computer storage devices. Along with, the input value can be processed using a conventional computer processing device. Similarly, in the sequential processing embodiment of the present invention, from the input value received during the time trial. The output value is calculated. In the sequential method, the processing device sequentially inputs the input value from the input vector. Acts to receive. Unlike the parallel method, the elements of the join weight matrix are data. It is sequentially stored in the storage device as a sequence. Sequential processing equipment is used as an element of the coupling weight matrix during the time trial. Acts to give pass-through output based on and join weight mat during time trial Acts to update the elements of the lix. Coupling weights as elements of a two-dimensional array Unlike the conventional method of arithmetic processing, in the sequential method, each element of the join weight matrix Is processed quickly with a specially designed sequence. Matori in the conventional method Because multiplication games are generally nested access loops (for raw and for columns) Concurrency is slower than the sequential method of the present invention. The present invention also provides a system for updating the binding weight matrix during an early trial. Served. Enter the system to update the join weight matrix during a time trial The relationship between a processing device that operates to receive a value from a force feature vector and a feature variable Contains a storage device that contains a binding weight element to identify. The processing device receives the input Acts to update the join weight element based on the non-disappearing value of the vector. Others Unlike equations, the processing equipment of the present invention learns differently for each input vector it receives. Created to update join weight elements based on component learning weights that are weights Move. Exact relations between feature variables by using component learning weights You can ask for a person in charge. Further, in both the parallel processing embodiment and the sequential processing embodiment of the present invention, information is provided. The output value and learning value are evaluated and co-rated by the controller unit in the processing system. Can be controlled. The relative effect of each input vector on previous learning Learning weights that automatically adjust the learning weights for each trial as they are generally adjusted A controller can be provided. In addition, the user interacts with the system Gives a desired learning weight that is different from the learning weight automatically provided by the system. Can be obtained. Also, the system will first receive the measurements received by the system. Features that act to convert to input feature vectors used for pass-through and learning A function controller is provided by the present invention. Characteristic function controller is default The system either gives the initial join weight or receives the join weight element from the outside. Operates so that the user can give an initial weight if desired. In addition, the learning weight controller complies when anomalous deviations in input values occur. The learning function of the data system can be disabled. In addition, the characteristic function controller The trawler is 1 of the current measurement and the corresponding measurement stored from the previous trial. It operates to generate various statistics such as next difference and identify sudden changes in measured values. Sudden changes in the input value may indicate that the instrument receiving the input value is defective. To. In addition to the physical examples of the present invention, some processes have been carried out by the present invention. To. The process of the present invention involves input vector during a time trial in a processor. Receives m [IN] (f) and enters based on join weight element during time trial The output value is calculated from the disappearance input value of the force vector, and the input vector is input during the time trial. It involves updating the join weight element based on the input value of the file. The process of the present invention is further combined based on the component learning weights described above. It can include a step to update the weight element. The learning weight element is global Learning that receives training weight 1 and is an indicator of the learning weight before each input vector Viability that receives the history parameter λ and indicates the degree of disappearance of the input feature vector (viability) Receives the vector ν (f) and multiplies these values together Get the learning component weights (ie l (C) (f) = lν (f) l (f) )) Can be calculated. The present invention allows the binding weight matrix to be updated quickly during the same trial. The system combines the average vector μ [OUT] of all the received feature vectors. Pass on a value by using it as part of the weight update process. Pre-input Use the mean vector to calculate various output values and system parameters As a result, the update can be performed quickly. Pre-mean vector μ [IN] is before Is equal to μ [OUT] from the measurement trial of. The process is the first trial If so, μ [IN] is for the system default value, preferably the user-supplied value. Can be equal to the value 1.0.
The element of [OUT] is the following process formula Is calculated by.<img file="JP2000511307A_D0001.tif" /> In the process of the present invention, the intermediate pass-through vector e [IN] is used to set the join weight element. It also includes updating.
[IN] Vector elements can be calculated according to the following equation Can be done.<img file="JP2000511307A_D0002.tif" /> The join weight matrix can be updated using the following process formula.<img file="JP2000511307A_D0003.tif" />Here,<img file="JP2000511307A_D0004.tif" />And d = e [IN] ω [IN] e [IN]<sup>T</sup>= xe [IN]<sup>T</sup><sup></sup>In the update process, ω [IN] = ω [OUT] from the previous trial. If the current trial is the first trial, ω [IN] is the system default It can be equal to the identity value, preferably the identity matrix or the user-provided value. Wear. Elements of the output vector m [OUT] passed through according to the following equation during the pass-through process Is calculated.
[OUT] (f) = μ [IN] (f) + e [IN] (f) (2-ν (f)) + x (f) (ν (f) -1) / ω (f, f) Other values used by the process of the present invention will be described later. In addition, the present invention provides a large number of processor pairs for computing x vectors. Access methods are provided. This process is uniquely paired during the time trial Access a large number of processors and connect paired processor units Search for each join weight element located at the following switching junction and each pro E [IN] (f) located at Sessa was connected to the switching junction Transfer to the other processor, then all processor pairs in the system set the corresponding value of x It includes calculating the moving sum of e [IN] ω [IN] until it is calculated. Access each set of processors that updates the join weight element by the process of the invention A way to do it is also provided. This process is uniquely paired during the time trial Located at a switching junction with access to a large number of processors Coupling weight element by one processor located at the switching junction The search weight element is updated by the processor that searched and searched for the join weight element. Transferring the new search weight element to the memory element at the switching junction And are included. Therefore, an information processing system that performs accurate learning based on the received input value It is an object of the present invention to provide. It is the present invention to convert an input measurement value into an input feature value during a single time trial. Is another purpose of. Learning and performance (measurement and feature conversion) during a single time trial It is another object of the present invention to perform a bride). Another object of the present invention is to pass on the vanishing value from the non-disappearing value. Another object of the present invention is to identify anomalous input feature deviations. A system for rapid learning and performance during a single time trial Is another object of the present invention. Another object of the present invention is to identify sudden changes in input feature values. Another aspect of the present invention is to provide a system for rapidly parallel processing input feature values. Is the purpose of. Another aspect of the present invention is to provide a system for rapidly and sequentially processing input feature values. Is the purpose of. Provides a system in which a large number of parallel processing devices can communicate between each processing device in the system. Is another object of the present invention. To provide a system capable of accessing a large number of parallel processing devices as a pair. That is another object of the present invention. Communicating between paired processors in a minimum number of steps is another aspect of the invention. One purpose. It is another object of the present invention to provide a process for achieving the above object. .. These and other purposes and advantages of the present invention are described below with reference to the accompanying figures. It is obvious if you read the Ming dynasty. A brief description of the drawing FIG. 1 is a diagram showing an embodiment of the present invention. FIG. 2 is a block diagram showing an embodiment of a parallel processor according to an embodiment of the present invention. FIG. 3 is a block diagram showing an example of a sequential computer according to an embodiment of the present invention. FIG. 4 shows a pixel value array that can be arithmetically processed according to an embodiment of the present invention. The figure which shows. FIG. 5 shows the joint access memory and the pro used in the parallel embodiment of the present invention. Sessa circuit layout. FIG. 6a shows a node switch in the joint access memory of the embodiment of the present invention. The figure which shows the details. FIG. 6b shows a node switch in the joint access memory of the embodiment of the present invention. Side view of the details. Fig. 7 shows the joint operation during the intermediate matrix / vector calculation processing of the parallel embodiment. Timing diagram of Seth memory control. FIG. 8 shows a switching junk of the joint access memory according to the embodiment of the present invention. Joint access memory control timing for update operations related to operations Timing diagram. FIG. 9 is a diagram showing processing time interval adjustment of the parallel embodiment of the embodiment of the present invention. FIG. 10 shows a screen of the entire system implemented as a parallel embodiment of the embodiment of the present invention. Figure. FIG. 11 shows a blow of the entire system implemented as a sequential example of the examples of the present invention. Figure. FIG. 12 shows the communication connection of the controller used in the parallel embodiment of the embodiment of the present invention. Is shown. FIG. 13 shows the communication of another controller used in the sequential embodiment of the embodiment of the present invention. Indicates a connection. 14 to 22 show the steps of the process performed in the examples of the present invention. Flow diagram. Detailed explanation Operation overview Then refer to the drawings where the same numbers represent the same parts throughout several drawings. Simultaneous learning and performance information processing (CIP) made according to Ming's example Shows a neurocomputing system. CIP system with reference to Figure 1. 10 is realized by the computer 12 connected to the display monitor 14. ing. Computer 12 of CIP system 10 is a data acquisition device (DAD) 1 Receives evaluation data from 5, which is numerous at various times via connecting line 16. Measurements can be provided. Data acquisition computer board and related software Data acquisition devices such as software are commercially available from companies such as National Instruments. Has been done. Computer 12 also from the conventional keyboard 17 via input line 18. It can receive input data and / or operating specifications. One pair at some point Receiving and responding to an input measurement of is called a trial here, and a set of inputs The value is called a measurement record. Generally, when the CIP system 10 receives an input measurement record, the system Determines (learns) the relationships that exist between the measurements received during the trial. Tiger When the measured value in the ear disappears, the CIP system 10 becomes the non-disappeared current measured value. Both give the expected pass-through values based on the previously learned relationships between the previous measurements. The CIP system 10 receives the measurement record from the data acquisition device 15 and measures the measured value. Is converted to a feature value. Necessary for learning, that is, passing on, by converting measured values to feature values The number of learning parameters is reduced. Some measurements have feature values and other values calculated from them The constant value has disappeared or the monitored measurement of the system is different from the previous measurement A useful day for predicting or passing on values when always checking to see if they are off. Provide data. Upon receiving the input measurement record at the start of the trial, the CIP system 10 will each As fast as possible when the input record arrives (ie system 10 is real at the same time) Perform the following operations. Subtract the input feature value from the incoming measurement Extraction (simultaneous data reduction), identify abnormal input feature values or trends (simultaneous) Monitoring), estimating the disappearance feature value (that is, passing it on) (simultaneous judgment), and learning the feature And update the mean learning feature variance and learning interconnect weights between features (simultaneous learning) .. The CIP system is (a) instrument monitoring in a chemical or radioactive environment, (b) guard Star-mounted measurement monitoring, (c) Tracking missiles during unexpected attacks, (d) Inpatient care supervisor Continuous and adaptive, such as vision and (e) forecasting and monitoring of competitors' pricing strategies It is useful for many applications such as. Being fast in one application is another application Less important than. As a result, the CIP system is traditional (ie, sequential). An example of an computer or an example of high-speed parallel hardware is provided. Speed is not a major concern in some applications, but CIP's high speed is a broad utility It is advantageous for Liti. Sequential CIP examples are traditional statistical for two reasons It's faster than anything else. First, the CIP system is average rather than offline learning Using row updates, secondly the CIP system first covariates like traditional statistical methods Rather than computing the matrix and then inverting it, the inverse element of a covariant matrix Update directly. Parallel matrix reverse update allows fast CIP implementation Is done. When implemented using a sequential process, the CIP response time is the data used It increases as the square of the number of features increases. However, using parallel processes The CIP response time will increase as the number of features used increases. I'm sorry. In a parallel system, a processor is provided for each feature. resulting in , The response time of parallel CIP is faster than the response time of sequential CIP by several times the features used. Become. Overview of parallel system A parallel example of the basic subsystems of CIP system 10 with reference to FIG. Shown. Before examining the subsystem in detail, refer to Fig. 2 for arithmetic CIP. Review the overview of. CIP subsystem is system bus 19, transducer Includes 20, kernel 21 and manager 22. Transducer 20 And kernel 21 to perform the various simultaneous operations mentioned above It works continuously. Input measurements are first input by transducer 20 Is converted to. The input features are then processed by kernel 21 and passed on (ie) Output) Features, update learning parameters and monitoring statistics are obtained. Next, the output feature value is It is converted into a pass-through (ie output) measurement by the transducer 20. Mimic Jar 22 performs parallel operation of Transducer 20 and Kernel 21 Adjust and refine system operations from time to time. The basic components of Transducer 20 are recently featured Memory (RFM) 25 And an input processor 24 having an output processor 26. Input processor 2 4 and output processor 26 are input and output control units, respectively. Controlled by 27 and 28. Recently feature memory 25 is the previous tiger Stores a preselected number of input feature values m [IN] obtained from the ear. (Here all vectors are low vectors). As mentioned above, the most stored Simultaneous entry to current trial using near features with input measurements j [IN] The force feature value m [IN] can be calculated. At the beginning of each trial, Input Pro Sessa 24 takes the input measurement vector j [IN] and the corresponding ambiguity vector p. Receive. The ambiguity vector element makes the input measurement vector element non-disappearing or disappearing Identify as. Next, the input processor 24 changes (a) the input vector j [IN] into a certain input feature. These converted input features have recently been converted to other converted features in feature memory 25. In combination with the input feature vector m [IN], and (b) the ambiguity vector Convert p to the corresponding viability vector ν. Like the ambiguity vector Viability vector elements identify input vector elements as non-disappearing and disappearing To do. At the end of each trial, the output processor 26 has an output feature vector m [OU Receive T]. Output processor 26 corresponds to output feature vector m [OUT] Output Measured value Converted to vector j [OUT]. Input measurement vector j [IN received by transducer input processor 24 ] Contains the value of the input measurement vector from DAD15 (see Figure 1). D Given externally by AD15 or internally by manager 22 The ambiguity value that can be obtained is whether the ambiguity value of the measured value is 0 (disappearing) or 1 (non-erasing). Loss) or some intermediate value (combination of disappearance and non-disappearance values described below) Indicates whether. j [IN, as determined by the corresponding element of p being 0 ] Element disappears, then the non-disappearing element of j [IN] and / or learns before The corresponding element of j [OUT] is passed on based on the obtained information. The pass-through process Conversion from measurements to features within the transducer input processor 24, followed by Passing on the vanishing feature values in the kernel, followed by the transducer output processor The conversion from the pass-through feature in the service 26 to the pass-through measurement value is used. Transducer input processor 24 is a manager before simultaneous operation -Calculate feature and viability values according to the function determined by 22 ..
Each feature element in [IN] is a function of the measured value element in j [IN], and each of ν The viability element is the corresponding function of the ambiguity element in p. For example, m [I The first characteristic function in N] is the sum, m [IN] (1) = j [IN] (1) + j [IN ] (2), and the second characteristic function in m [IN] is the product, m [IN] (2) = j [IN] (1) j [IN] (3). Each feature buy The ability value is the product of the ambiguity values of the measured values, which are independent variables in the characteristic function. .. For example, the ambiguity values of the above three measurement functions are p (1) = 1.0, p (2) = If 0.5 and p (3) = 0.0, then the above two feature viability values are ν (1) = 1.0x0.5 = 0.5 and ν (2) = 0.5x0.0 = 0.0 Become. The feature viability element is calculated as the product of the corresponding measured ambiguity elements. Each CIP system input measurement is treated as the average of non-disappearance measurements from the larger set And some of them may have disappeared. Corresponding ambiguity of each input measurement The value is treated as a percentage of the amount of non-disappearing components within the large set. probability From theory, addition or product synthesis characteristic functions depend on some such input measurements. If the distribution of disappearances is independent between the measured values, then all quantity measurements are non-disappearing. The expected ratio of the composition function term is the product of the component measurement ambiguity values. CI Because the feature viability values in the P system have an interpretation of this expected ratio , The feature viability value is calculated as the product of the component measurement ambiguity values. Input measurements j [IN] and ambiguity values p are transducer input processors 2 After being converted to the input feature vector m [IN] and the viability value ν by 4. Then kernel 21 starts the next in-trial operation. Kernel 21 For input to, the resulting feature value in m [IN] and the corresponding viabi in ν Contains a lativity value and an input learning weight of 1. Kernel 21 is a feature 1 processor Sa 31<sub>1</sub>From feature F processor 32<sub>F</sub>Up to kernel control module 3 2, and buses 45-45<sub>F</sub>From processor 31 to processor 31 via<sub>F</sub>Contact Contains the continuous joint access memory (JAM) 23. kernel The output from 21 is the passed-through feature value in m [OUT], connection 41<sub>1</sub>From 42<sub>F</sub>O And 41<sub>JAM</sub>Characteristic function monitoring statistics sent to the manager via, and connection 40<sub>1</sub><sub></sub>From 40<sub>F</sub>Includes feature value monitoring statistics sent to the manager via. Preferably kernel processor 33<sub>1</sub>From 31<sub>F</sub>Is the processor 31<sub>1</sub>From process Ssa 31<sub>F</sub>Perform basic arithmetic functions to reduce the cost and size of Use an arithmetic unit (ALU). As those skilled in the art know, such The basic processor is Mentor, a product of Menthol Graphics.<img file="JP2000511307A_D0005.tif" />Can be designed. Kernel processor 31<sub>1</sub>From 31<sub>F</sub>Is a non-disappearing element of m [IN] and / if Or pass on the vanishing feature values based on the previously learned kernel 21 parameters, and each Learned parameters residing in the roseser and joint access memory 23 Updates to produce the monitoring statistics used by Manager 22. Postscript Kernel processor 31 to do<sub>1</sub>From 31<sub>F</sub>Is the 2nd mode of inter-processor communication Use the tep to transfer the relevant values from each processor to each other processor. The kernel processor also calculates the distance measure d within the kernel distance ALU34. Communication between distance ALU34 and each kernel processor is connected 35<sub>1</sub>From 35<sub>F</sub>Through Will be done. Kernel input learning weight 1 is non-negative-input ambiguity value and viability Like values-quantity / probability measure. Learning weight 1 for each trial is CIP Processed by the stem as a percentage of the quantity count, the molecule is in the current trial It is the number of quantification vectors to the denominator, and its denominator is the total number of quantification measurements used in the previous training. It is a total. Therefore, the current input feature vector m [IN] has a high learning weight of 1 value. Has more learning parameter updates than having a lower learning weight 1 value The impact is large, which results in the input feature vector m [IN]. This is because it contains a high proportion of the sum of the ambiguity measurements. Usually the learning weight 1 is It is given as an input variable during each trial, but the learning weights are optional as described below. CIP system manager 32 can also occur. Kernel pass-through, memory update and monitoring operations are additive for non-lost feature values It is based on a statistical regression framework that predicts vanishing features as a function. Times Within the return framework, the weight of passing on each vanishing feature value from everything else is well known. ing. The weight formulation used for pass-through is a function of the inverse of the sample covariant matrix. To. In the traditional regression method, the FxF covariant matrix is first based on the training sample. Is calculated, then the covariant matrix is inverted and then the regression weight is calculated as the inverse function. It is calculated. The traditional method uses all measurements received up to the current input trial. It includes storing the including learning set and performing arithmetic processing. Get the inverse matrix All previous measurements are needed to first calculate the current covariates that can be Therefore, a conventional system typically stores all measurements. Unlike traditional statistical operations, the CIP kernel 21 operation Is (a) the inverse element of ν and other parameters learned by its trial , (B) Based only on the incoming feature value m [IN], and (c) Input learning weight 1 And directly update the inverse element of ν. Therefore, the CIP operation is a learning data set. Arrives quickly without the need to store and compute or invert the covariant matrix Su Information and pace can be maintained. The process of updating the inverse element of ν is for the traditional learning of CIP. C Traditional learning from offline learning data due to high-speed update capability for each IP trial Statistically appropriate fast improvements to Xi are made. As a result, the CIP system 10 improves the prior art. Continuing with reference to Figure 2, the joint access memory 23 is a possible Fx. (F-1) / 2 Each pair of features contains one feature interconnect weight. Special Corresponds to the lower triangular element of the inverse element of the symptom binding weight ν. ν The main diagonal element of the inverse element is also Kane Used during pass-through and feature function monitoring, and modified during parallel learning. ν Inverse element pair The individual elements of the angle set reside in their corresponding kernel processors. When kernel 21 passes on the feature value in m [OUT], the passed-through feature vector Send m [OUT] back to the transducer output processor 24 via line 33 and Therefore, the passed-through feature value in m [OUT] changes to the passed-through measured value in j [OUT]. It is converted and output to the system by the transducer output processor 26. A certain mo In Delling situations, only simple output conversion is needed. For example, CIP system Output processor 2 if the feature contains the original measurement along with the product function of the original measurement 6 decides to exclude everything except the passed-through measurement set from the passed-through feature set Convert more passed-through features into passed-through measurements. In another modeling situation, more Elaborate conversions are available. For example, the CIP system feature is Trans A set of measurements can be converted to their mean during the user input process, in which case The transducer output processor 26 passes all the measured values passed on to them. Set to the average value. If the passed-through measured value j [OUT] is generated as an output, (a) the instrument fails. (B) Econometric forecasting operation period that directly replaces the measured value during the period, etc. Predict measured values before making measurements, etc., (c) Potential failure product classification operation In several ways, including predicting measurements that should not occur during the period, etc. The output will be useful. Manager 22 monitors and controls CIP system operations .. Manager 22 subsystem has a CIP system user interface Proposal Coordinator 38 to serve and executor to direct overall system control Instead of the l value supplied externally from the device 39 and the data acquisition device 15 (Fig. 1). A learning weight controller 40 that gives 1 to kernel 21 and a measurement-feature function function Includes feature function controller 41 to establish and modify. Same as CIP system In time operation, kernel and transducer 20 are in simultaneous mode Is active in. Carne when the system is operating in simultaneous mode Le 21 and Transducer 20 are set by Executive 48 Input measurements, ambiguities and training weights according to system control parameters Operates continuously based on. Input measurement specifications and feature calculations for these parameters Includes specifications, inter-module buffering specifications and output measurement specifications. Parallel movement During the work, the CIP system has passed on feature values, feature value monitoring statistics, and updated learning cars. Produces flannel parameters and input output measurements. CIP system includes kernel 21, learning weight controller 40 and coordinator It is also possible to carry out the feature value monitoring operation performed by the data 49. Feature value monitoring During operation, each kernel processor 31<sub>1</sub>From 31<sub>F</sub>Connect monitoring statistics 40<sub>1</sub>From 4 0<sub>F</sub>Send to manager 22 via. Learning within Manager 22 during each trial The weight controller 40 uses deviation monitoring statistics to (a) prior to that feature. Passed on when the mean value calculated from learning and (b) features disappeared Evaluate the unexpected degree of each feature with respect to the feature value. Is it each kernel processor? The feature value statistics sent from are for observations, learnings, regressions and processor features. Contains the learning variance value. Learning weight controller 40 uses feature value monitoring statistics And calculate the parallel feature variance measure. These variance measures are then learned weight control From La 40 to Coordinator 38, surveillance graphics are created and Sith It is sent to monitor 14 via tembus 19. Specify feature pass-through, feature value monitoring, and learning parameter update operations at the same time Besides, CIP Manager 22 sometimes specifies feature function evaluations and assignments. Rui or CIP manager 22 controls learning and assignment. Features Function Evaluation and allocation is connection 41<sub>1</sub>From 41<sub>F</sub>Together with other weights in processor 1 via In parallel port 41<sub>JAM</sub>Mutual coupling in joint access memory 23 via Feature function control in manager 22 by accessing weights at the same time Implemented by La 41. Features Function controller 41 first adjusts interconnect weights Provides all redundant or unnecessary features, that is, useful information for learning and pass-through Identify no features. Features Function controller 40 is then the transducer input controller Command the Rossessa 24 to combine redundant features via control line 43 and fail Remove key features or add new features. Similar to CIP pass-through and learning parameter update operations, CIP function supervisor Visual and control operations are based on a statistical regression framework There is. For example, the partial correlation coefficient needed to identify redundant or unwanted features And are all the multiple correlation coefficients ν inverse elements resident in the joint access memory 23? Can be calculated from the kernel to perform such sophisticated operations 21 Architectures that are very similar to architectures can be used. Sophisticated operations are as fast as parallel kernel 21 operations You can't, but with the parallel refinement processor described below, the kernel in progress It can be performed in harmony with 21 operations at almost the same speed. Calculate the learning weight schedule based on the probability / quantity basis of the learning weight interpretation, (a) Each input feature vector has the same overall impact on parameter learning, etc. Impact learning, (b) More recent input features Input features less recent than vectors Conservative learning in which vectors have a higher overall impact on parameter learning Rivera with a lower overall impact on recent input feature vectors than on, and (c) Le learning can occur. Using the learning weight controller 40 If you give learning weights to the CIP system in a basic format, the system is the same It is programmed to give only impact learning weights. In another form, learning Weight controller 40 uses CIP system monitoring statistics to pass on anomalous accuracy The lend can be identified. If the pass-through accuracy drops sharply, learn in advance For a new set of situations where the impact on the parameters you have Based on the assumption that the pass-through accuracy will be lower, the learning weight controller 40 will learn. Change the training weight calculation schedule to create more liberal learning. Feature value monitoring If the measurement behavior is strange, the learning weight modifies the element of the ambiguity vector p. When You can also. Overview of conventional sequential computer systems CIP of a conventional computer with one central processing unit, referring to Figure 3. An example of the system 10 is shown in a block diagram. Sequential CIP system 11 conventional one-by-one Next Computer As a basic component of Example 12a, a transducer input 24a process, kernel process 21a, transducer output process 26 a, coordinator 49a, executive 39a, learning weight controller 40 a, and feature function controller 41a is included. Sequential CIP system 11 Performs the same basic functions as the parallel CIP system 10. However, the conventional The computer system uses only one central processor. Therefore , Kernel 31a input because it uses only one processor for operation The data processing time is longer than that of the parallel CIP calculation system 12. Similar to the parallel system embodiment, the sequential system has the input vector j [IN] and Receives the ambiguity vector p. Input vector j [IN], similar to parallel system Is also converted to the input feature value m [IN]. As mentioned above, the ambiguity value p is Converted to earability ν. Kernel process 31a features feature input value m [IN] Receives a viability value and a learning weight of 1 from the system. Kernel 21a The process is stored in conventional memory 301 distributed by executive 39a Create an output feature vector m [OUT] based on the combined weights obtained. CIP Sith As described above for the system 10, the output feature vector m [OUT] is the output transistor. Transferred to the device 26a and changed to the output measurement value j [OUT] for external use. Will be replaced. Executive block 39a represents the main function of the computer sequentially and other blocks Represents a CIP subroutine. The memory 301 of the kernel subroutine embodiment is the same. It has a conventional data array format that is well known to vendors, and all shared memory is the main executor. It is distributed and maintained by the Tib function. Executive 39a first Cody Initialize the CIP system by calling Nator 39a, then it measures Specify the system specifications provided by the user, such as the vector j [IN] and the number of feature functions. -Obtained via board 18. Then Executive 39a learns accordingly Distribute meter memory and other storage devices. During each parallel trial, the Executive 39a program is a transducer Call the input processor 24a subroutine, followed by the kernel 21a subroutine Call Chin, then call the transducer output 26a subroutine. If Executive Program 39a is set to do so first, then Executive 39a is a learning weight controller 40a Sable at the beginning of each trial Call Chin to receive an input learning weight of 1 and the executive will be on each trial Finally, the feature monitoring statistics are given to the coordinator 38a and graphed on the monitor 17. It can also be displayed. For each trial of the sequential example, as in the parallel example Reads the input measurement vector j [IN] and the ambiguity vector p, followed by It involves writing the married measurement vector j [OUT]. But what In the conventional calculation, input file 15 and output file are used for input / output operations. Il 17 is used. The input file is a storage medium that receives the input value from DAD15. Read from the body. The output file will be CIP system in any way the user chooses. Used outside the system. In addition to parallel operations, as described above for parallel systems, it is carried out sequentially. Examples can sometimes take advantage of sophisticated operations. Initial in sequential version Executive 38a sometimes goes into parallel operation when specified by the user during conversion interrupt. During each such interrupt, Executive 39a features a feature control Call roller 41a, which receives the join weight matrix as its 1 input To. Next, the feature function controller 41a is redundant using the join weight matrix. And identify unwanted features, then return new feature specifications to Executive 38a. Next Executive 38a introduces new specifications during subsequent parallel operations Transducer Input Subroutine 24a and Transducer Output Subroutine Carry to Chin 26a. Operation and implementation details An example of CIP pass-through The following example shows some CIP operations. Refer to Fig. 4, three Three binary pixels that can represent different CIP input measurements vectors Indicates a. Each array has 9 measurement variables, x (1,1) to x (3,3) There is. The black square can represent the input binary value 1 and the white square represents the binary input value 0. Can be done. Therefore, the three arrays are (1,0,0,0,1,0,0, 0,1), (0,0,1,1,1,0,0,0,0) and (1,0,1,0, Can be represented as a CIP binary measurement vector j [B] with a value of 0,0,0) it can. As mentioned above, the CIP system uses each ambiguity value in p to j [IN]. Establish a role of disappearance or non-elimination of that corresponding measurement within. Therefore, If the values of all nine ps corresponding to the measurements in Fig. 4 are 1, then the nine of j [IN] All values are used for learning. However, if the two p elements are 0, then the two Indicates that the corresponding j [IN] value has disappeared and has a corresponding p-value of 1 From the seven j [IN] values of, two corresponding j [OUT] values are passed on. Continue to refer to Fig. 4 and p = (1,1,0, in the 101st trial. 1,1,1,1,1,0) and j [IN] = (1,0,?, 0,1,0,0,0 ,?), And the two ? Symbols represent unknowns. In this example, above The two pixels on the left and middle are black, and the pixels on the top right and bottom right are missing. The other 5 pixels are white. Therefore, the seven non-disappearing values in j [IN] The training is updated using and the result is j [OUT] = (1,0, j [OUT) ] (3), 0,1,0,0,0, j [OUT] (9)) and j [OUT] ( 3) and j [OUT] (9) are regression solutions based on previously learned parameter values. Using analysis, it is passed on from the other seven known numbers. Therefore, the previously learned pa The output pattern is 70a or 70c, depending on the parameter value. If the CIP system is configured for equal impact learning operations Between 70a and 70b patterns that can occur in previous trials 1-100 The most generated pattern in is passed on. For example, previous trials 1 to 100 In each of these trials, the nine values of p are all 1, and this is it. For una 40,19 and 41 trials j [IN] values are type 70a,7 It shall correspond to 0b and 70c. In this example, the CIP system is just another Because it is taught to anticipate 70c slightly more than one possibility 70a , Unknown top right and bottom right pixels in trial 101 harmonize with 70c, respectively Then, it is passed on as white and black. Transducer input operation Input measurements can be called arithmetic, binary and categorical. Altitude, temperature And arithmetic measurements such as time of day are via arithmetic operations such as addition, subtraction, multiplication and division. And have quotas that can be used in order. What happens to the measurements If you only get two states, the reading can be called binary and generally even one. It is represented by a value of 0. Non-arithmetic measurements with 3 or more states are categorized Can be called. Therefore, the CIP measurement vector j is arithmetic, binary and mosquitoes. It is classified into tegory subvectors (j [A], j [B], j [C]), and each sub vector The couture can be of any length. Depending on how the option is specified, the CIP system will (a) arithmetic And binary measurements converted to feature values on the transducer input processor 24 Or (b) directly the CIP kernel without converting arithmetic or binary measurements Send to. In contrast, the CIP system converts categorical measurements to equivalent binary features. Convert. To represent all possible contingencies between categorical variables The stem has one of the possible C values, each category measurement j [C], To a binary feature vector m [C] with a C-1 element with a possible C value Convert. For example, the possible values of a categorical variable are 1,2,3 and 4 If, then the resulting category feature vector has the corresponding values (1,0,0), (0) , 1,0) and (0,0,0). After the input processor 24 converts the category measurements to binary features, the transformer Jucer is an arithmetic measurement in the original input format, a binary measurement in the original input format, or a category. It contains only features that are binary equivalents of Golly measurements. These are all arithmetic features Processed as and sent directly to the kernel, or by the input processor 24 It can be converted to other arithmetic features at will. It s not limited, but it s like this. Arithmetic measurements, arithmetic measurements, binary measurements or Secondary and higher cross products between arithmetic and binary measurements, the mean of such features , Main component features, orthogonal polynomial features, and any such features Re Combined with such other features within RFM25 calculated from recent measurements Contains composite values. Recent features are functions of pre-observed feature values stored in the recent feature memory 25. Can be used to monitor and pass on (or predict) parallel feature values as Wear. For example, monitoring outliers in measurements from a scientific process to indicate a system failure It shall identify sudden changes in measurements that may occur. CIP system, each Each is stored in the measurements and recent feature memory 25 for parallel trials. Generates a CIP first-order difference feature, which is the difference from the same measurement for the previous trial By identifying sudden changes. CIP by generating first-order difference features The system can quickly learn the mean and variance of these features, and so on. This allows the CIP system to identify outliers in the first-order difference feature as a display of sudden changes. be able to. As a second example of using feature memory recently, the present of continuous processes Value prediction can be used to predict parallel values before actually observing them. CIP The system recently used feature memory 25 for each parallel as well as the last 5 measurements Features can be generated as parallel measurements. Then the CIP system is the first 6 How from the last 5 to the 6th during Trial 6 using one observation Learn whether to pass on the value, then the CIP system when the 7th value is missing Pass on the 2nd to 6th to 7th values at the beginning of Trial 7 Can then end Trial 7 after receiving the 7th value that has not disappeared The learned parameters can be updated, and then the CIP system is 8th. Is it the 3rd to 7th value at the beginning of Trial 8 when the 3rd value is gone? The eighth value can be passed on, and so on. Ambiguity and viability details In addition to operating with binary ambiguity values as discussed for Figure 4, the CIP system Can be implemented to work with ambiguity and viability values between 0 and 1 it can. Values between 0 and 1 are for some situations such as preprocessing outside the CIP system. It occurs naturally. For example, the average of one or two rows of an array instead of nine measurements Only three measurements corresponding to the values are given to the CIP system based on the data in Figure 4. It shall be. In this example, the CIP system is (a) its three components If all of the pixels are non-disappearing, set the ambiguity value to the mean value to 1 and ( b) Set the ambiguity value to 0 if all 3 pixels are missing, and (c) 3 If one or two of them disappear, set the ambiguity to a certain intermediate value. One If the value of is lost, then the appropriate measurement is the mean of the non-missing pixel values, which is appropriate. Ambiguity value is the ratio of non-disappearing pixel values to the total number of possible pixel values Is. The user of the CIP system does not calculate the ambiguity value objectively, but rather the confidence of the measured value. You can also use an ambiguity value between 0 and 1 if you want to evaluate the reliability subjectively. To. CIP system for parallel values of other measurements as well as values prior to that measurement The system treats the ambiguity value between 0 and 1 as a weight for learning the measured value. In ambiguity The process and formulation of the based weighting scheme will be described later. Kernel learning operation As mentioned above, the CIP system provides useful data based on a set of measurements. To serve, identify related parameters and learn the relationships between many input measurements Carry out the exact process to do so. CIP kernel guesses learned parameters By updating the fixed value, it is learned during each trial, and the parameter estimated value is special. Vector μ of symptom mean, matrix ω of coupling weight, vector ν [D] of variance estimate ] (Diagonal elements of the variance-covariant matrix ν described above) and the learning history parameters Quant λ is included. The update formula for learning parameters will be described later, but it is basic. Check for a simple version of the first parameter to show its properties Defeat. If all leading and parallel viability values are 1, the average update formula is: It has a simple form.<img file="JP2000511307A_D0006.tif" />The μ [OUT] term represents the average of all previous measurements and includes parallel measurements. To. According to the above examination of the learning weight, the value of μ in Eq. (1) goes toward m [IN]. For a large value of 1, it changes more than for a small value of 1. However, (1 ) Exactly reverses the history of different ambiguities for different elements of μ [IN] Equation (1) is preferably modified because it may not be reflected. Instead, (1) The formula is an element of the learning history parameter λ that tracks the previous learning history at the future element level. Therefore, it is modified to combine μ [IN] with m [IN]. Equation (1) is justified in the CIP system's quantitative framework as follows: Pulled out. The learning weight 1 is the parallel quantity count of m [IN] and the CIP system. The ratio to the parallel quantity count q [PRIOR] related to μ [IN] And then μ [IN] is the mean of q [IN] pre-quantity counts and m [IN] ] Is the mean value of the q [PRIOR] parallel quantity count, depending on the algebra Equation (1) includes all q [PRIOR] parallel counts and all q [IN] pre-counts. It can be shown to be the overall mean value based on the und. Equation (1) is applied only when all the viability values are 1.
[0 Individual elements of μ [OUT] according to the various viability histories of the UT] element The CIP system is preferably more precise than Eq. (1) in order to properly weight Use various functions. The average update equation for any element μ (f) of μ is expressed by the following equation. To.<img file="JP2000511307A_D0007.tif" />(f = 1, ..., F-such f labels are used in the rest of this specification. Used to indicate a ray element). About all feature vectors during the trial A single learning weight 1 in equation (1) used is the feature vector for each component. Equation (2), except that it is replaced with a different learning weight 1 [C] (f) for the element. Is similar to equation (1). Therefore, each feature vector element is the same as in Eq. (1). It has individual weights rather than the same weights. These component learning weight Only depends on the parallel and pre-learning viability values of the form:<img file="JP2000511307A_D0008.tif" />After the learning history parameter λ has also been updated and used to update the feature mean, Keep a pre-learning execution record for each feature.<img file="JP2000511307A_D0009.tif" /> The remaining learning parameters ν [D] and ω are the covariant matrices ν and ν, respectively. It is an element of the inverse element. In addition, ν is not only a feature value generally known as an error, but also a special value. It depends on the deviation from the mean value of the levy. As a result, ν [D] and ω are feature values Can be updated as a function of error vectors of the form: it can. e = m [IN] -μ [OUT] (5) A suitable formula for updating the elements of ν can be:<img file="JP2000511307A_D0010.tif" />(The letter T in Eq. (6) used here indicates vector transpose). Equation (3) is the same Since it is based on the CIP quantity counting framework, the general form of equation (6) is (3). ) Is similar to the equation. Overall mean value μ [OUT of the quantity count value assumed by Eq. (3) ] Is generated, Eq. (6) is a square deviation from the mean vector μ [OUT]. And yields the overall mean ν [OUT] of the cross product value. An appropriate formula for updating the element of ω based on equation (6) is shown by the following equation.<img file="JP2000511307A_D0011.tif" />Here, d = eω [IN] e<sup>T</sup> (8) Eq. (7) has the form of Eq. (6) when the second term of Eq. (6) is known. It is based on a standard formula that updates the inverse element of the lix. Equations (7) are also equations (3) and (6) ) Is based on the same quantity counting principle as the equation. Equations (5) to (8) are preferred approximate versions of the CIP error and update equations. It's just that. However, various preferred alternatives have been used for four reasons. Is done. First, the CIP kernel learned more quickly using μ [IN]. Error vector formula (5) because the meter can be updated to facilitate high speed operation ) Corresponding CIP system is based on μ [IN] instead of μ [OUT] To. Second, in the CIP example equation preferred in equation (5), the corresponding parallel viabi If the degree ν (f) is less than 1, each element of the error vector e decreases toward 0. Will be done. Due to this decrease, the corresponding element of m [IN] of the e-vector is low. For updating the elements of ν [D] and ω for each element of the e-vector, even if it is a lit value And properly given a small role. Third, the CIP kernel requires all elements of ν It is not necessary to use the element of ν [D], and [D] represents the diagonal element. Finally, formula (6) And (7) is accurate because it calculates the previous μ [OUT] and ω [OUT] values. Where the previous μ value used for is the same as μ [OUT] for the current trial Only if. All such μ values change during Eq. (7), which is preferable CI. In P Example, the CIP system uses Eq. (7) after modifying it appropriately. Is it equation (5)? The preferred alternative to Eq. (8) is:<img file="JP2000511307A_D0012.tif" />and<img file="JP2000511307A_D0013.tif" />Here<img file="JP2000511307A_D0014.tif" />And d = e [IN] ω [IN] e [IN]<sup>T</sup>= xe [IN]<sup>T</sup> (14) The values of ω [IN], μ [IN] and ν [D, IN] are in the learning parameter memory. Represents the output value from the previous stored trial. Kernel pass-through operation CIP features pass-through equations are linear regression theory and CIP storage, velocity, input / output flexibility It is based on a formulation that fits the flexibility and parallel examples. All other features Regression weights that pass on arbitrary features from ω are available in ω, from the element of the ν inverse element. Efficient kernel operation is possible because it can be calculated easily. The CIP kernel passes each lost m [IN] element as a function of all non-lost elements And the vanishing and non-disappearing m [IN] elements are the corresponding buys of 0 and 1, respectively. Indicated by the ability ν element value. CIP application has 0 value, F features and via billi for trials containing all 1 values except the first element When using tee vectors, the CIP kernel is the first feature as all other functions. Pass only the levy. The regression equation that passes on the first element is shown in the following equation.
[OUT] (1) = μ (1)-{[m [IN] (2) -μ (2)] ω (2,1 ) + ... + [m [IN] (F) -μ (F)] ω (F, 1 )] / Ω (1,1) (15) Pass on other m [IN] elements as long as the m [IN] element to be passed on disappears The formula is the same as equation (15). The CIP kernel uses the improved equation in Eq. (15) to use any set of m [IN] elements. The CIP system can be operated even if the alignment is lost. Any element disappears If so, the kernel will use only the other m [IN] elements that have not disappeared. Pass on the disappearing m [IN] element. The kernel is also where m [IN] has not disappeared If so, always replace each m [OUT] element with the corresponding m [IN] element. CIP Sith The regression equation used by the system is also a parallel operation as well as efficient operation. Designed to do rations and make the most of other kernel calculations. Example For example, since e [IN] and x are also used for learning, the kernel uses Eq. (9) and And time by using the elements of e [IN] and x from equation (14) for pass-through And save storage. The kernel pass-through formula is expressed by the following formula.
[OUT] (f) = μ [IN] (f) + e [IN] (f) (2-ν (f)) + x (f) (ν (f) -1) / ω (f, f) (16) Monitoring operation The kernel has several controls for feature value monitoring and graphic display Create a meter. To do this, the learned feature mean vector μ [OUT] and feature variance ν [ D] and a well-known statistical monitoring measure called the Mahalanobis distance, (15) Contains d from the expression. The kernel also produces another set of regression feature values, which Is the passed-through value of each feature when the feature has disappeared. These regressions The value has the following format.<img file="JP2000511307A_D0015.tif" />(17) Given the above monitoring statistics from the kernel, the CIP system will be user-optimized. Statistics can be used in a variety of ways as specified by the application. One Applications include Mahalanobis distance d and standard squared deviations from the learned mean. To plot the deviation measure as a function of the number of trials, d [1] (f) = (m [IN] -μ [OUT] (f))<sup>2</sup>/ ν [D, OUT] (f) (18) The standard squared deviation value between regression values is expressed by the following equation.<img file="JP2000511307A_D0016.tif" />(19) Mahalanobis distance measure d and three deviation measurements of Eqs. (17), (18), (19) Degree is a powerful indicator of manual input behavior. Mahalanobis distance d is each observed All features because it is an increasing function of the square difference between the feature vector and the feature learning mean vector A useful global measure for vectors. The standard deviation measure d [1] (f) is A component feature for global measures to identify anomalous feature values Can help. The standard modulo measure d [2] (f) is a pre-learned mean. Which input features are from those regression values, not just based on other non-disappearing parallel feature values Indicates whether it deviates as follows. CIP systems also use special features in connection with their monitoring statistics Can produce useful information about symptomatic tendencies. For example, any interest The difference between the parallel feature value and the feature value from the immediately preceding trial for the feature. New features can be calculated. By the deviation measure obtained from Eq. (17) A useful measure of anomalous feature value changes is obtained. CIP system is also not a first-order difference Identify anomalous deviations from normal feature changes using a similar method based on quadratic differences can do. Therefore, the CIP system is a manifold outside the CIP system. Can give a variety of graphic deviation plots for dual user analysis To. The CIP system also uses deviation information internally to control learning weights. You can plan a symptom correction operation. For example, the system is global A preselected cutoff value for the distance measure d can be established. next The system sets a value of d that exceeds a preselected cutoff value for data entry device problems. Can be treated as evidence and the system will have future learning weights until the problem is determined Can be set to 0. Similarly, the component deviation measure d [1] (f) And after they exceed a pre-specified cutoff value using d [2] (f) You can set the measured value ambiguity or feature viability value to 0 with. Gaku By setting the training weight to zero, the input problem is the finest for future CIP operations. It is possible to prevent adverse effects on the degree. Distance measures for various measurement environments Because it follows a chi-square distribution, C is based on a stable statistical basis of the CIP system. The IP system will be useful for such judgment applications. As a result, known Kai The distance cutoff value can be inferred from the squared cumulative probability value. Learning weight control operation The CIP system uses a learning weight of 1 as part of system learning. Learning weight l is related to the parallel feature vector for the quantity count related to pre-parameter learning It is the ratio of the amount count to be done. From this base, the CIP system has an equal impact For learning weight sequences, i.e. equal number of quantity counts for each trial Create a sequence, based on. Learn weight sequence l (1), l (2), etc. When displayed with, the equal impact schedule has the following format.<img file="JP2000511307A_D0017.tif" />The constant R is a common quantity coun for all such trials for the initial quantity count. It is the ratio of The role of this ratio and the initial quantity count will be described later. In addition to providing equal impact learning weight sequences, CIP users and CIP systems The system can generate liberal or conservative sequences. liberal Sequence has more impact on the feature values of more recent trials, Conservative sequences have more impa to less recent trial feature values Give the act. For example, a learning sequence with all learning weights set to 1 A study that is liberal and has all learning weights set to 0 except the first learning weight. The learning sequence is conservative. Input CIP data is intermittent for liberal sequences Suitable and conservative sea when generated according to parameter values that change Kens is appropriate when more recent information is not as reliable as less recent information Is. Learning parameter initialization The CIP system uses the initial values for the learned regression parameters μ, ν [D]. Treat as if they were generated from the feature values observed by them. During the first trial The CIP kernel sets the initial value to the first value according to equations (2), (10) and (11). Combined with the information from the feature vector of to generate the updated parameter value, the second During the trial, the CIP kernel sets the initial value and the first trial value to the second vector. Combines with information from the torr to generate newly updated parameter values. This pro Seth repeats for subsequent trials. An initial para on overall learning after any trial with a positive learning weight of 1. The impact of the meter is smaller than the impact of the initial parameters before the trial I. As a result, as long as a very conservative learning weight sequence is not given to the kernel Therefore, the effect of certain initial regression parameter values is small after a small number of learning trials. It gets cold. Accurate pass-through by CIP system is necessary from the first trial The initial values of the learned regression parameters are important. Smell of such applications In order to perform accurate early pass-through, the CIP system uses keyboard as shown in Fig. 1. The user-provided initial regression parameter value can be received from the mode 17. The CIP system defaults to the learned regression parameters as follows: Give a value. The default value of each element of the mean vector μ is 0, and the join weight mat The default value of lix ω is the identity matrix, and each of the deviation vectors ν [D] is required. The raw default value is 1. Use the identity matrix as the initial default value for ω First, an initial pass-through feature value that does not depend on other feature values is created. Also early The identity matrix allows the CIP system to change feature values from the first trial You can marry. On the other hand, in the conventional statistical method, before any pass-through is performed. Requires at least a learning trial of F (F is the number of features). In addition to initializing the learned regression parameters, the CIP system has a learning history parameter. Initialize the elements of the data vector λ. The learning history parameter vector is the previous learning Instructs how much the input feature vector element affects learning. The default initial value for each element of the learning history vector λ is 1, which is the first Give the same learning impact to each input feature vector element during the learning trial of. Features Function monitoring operation The CIP system also monitors three features for graphic displays. Statistics can be carried out, squared feature multiple correlation χ [M] vector, tolerance band ratio value r And the partially correlated array χ [P].
Each element of [M] χ [M] (f) is of m A square-double correlation that passes on the corresponding feature vector element m (f) from another element. .. When performed voluntarily within the CIP system, the squared correlation follows well-known statistical properties. Can be interpreted. Such statistical properties are the minimum possible for the square correlation of features to be 0. It implies that if it is close to the maximum possible value of 1 instead of the Noh value, it can be predicted by other features. ing. Each squared multiple correlation χ [M] (f) is given by calculating the corresponding tolerance band ratio element r (f). It can also be used at will. Each element of r is as a ratio of two standard deviations Can be represented. The numerator standard deviation is the square root of ν [D] (f) and the denominator standard deviation<img file="JP2000511307A_D0018.tif" />Therefore, each r (f) passes on m (f) if all other m [IN] values have disappeared. Tolerance band width, whereas all other m [IN] elements have not disappeared The tolerance band width that passes on m (f). The partially correlated array χ [P] is a possible feature pair f and g (f = 1, ..., Contains partial correlation χ [P] (f, g) for F-1; g = 1, ... F) .. For each partial correlation, after adjusting the correlation with all other features, which is the correlation between the two features? It is an index showing whether it is only high. As a result, the user examines the partial correlation and elsewhere It can be determined whether any given feature is needed to pass on a given feature. The user also examines the rows of the partial correlation matrix and uses the pair of features separately. Instead, you can identify if you can combine them to produce an average. To. For example, two features are needed to pass on the third feature, as opposed to the first feature. It is assumed that each partial correlation is the same as the corresponding partial correlation for the second feature. .. Next, both of these feature values can be used as the third feature value without compromising the pass-through accuracy. It can be replaced with the mean value to be passed on. The benefits gained by the CIP system are the features sometimes made by Manager 22 Parallel operation capability linked with functional evaluation, parallel operation It will be carried out very quickly and the feature function evaluation operation will be carried out immediately. The CIP system can obtain the squared correlation value using the formula below, χ [M] (f) = 1-ν [D] [D] (f) / ω (f, f); (23) The system can use the following equation for the tolerance band value, r (f) = (1-χ [M] (f))<sup>1/2</sup>; (twenty four) The system can use the following equation for the partial correlation value.
[P] (f, g) =-ω (f, g) / (ω (f, f) ω (g, g))<sup>1/2</sup>(twenty five) In addition to providing the above-mentioned feature function evaluation statistics to the user, the CIP system provides join weights. The matrix ω can be supplied to the user for the user to modify and interpret. Example For example, does the user calculate the major component coefficients, orthogonal polynomial coefficients, etc. from ω? It can be evaluated to identify the essential features that meet the user's needs. Essence Once identified, the user may reformulate the input transducer function. Features can be supplied to the outside of the CIP system. Features Function control operation In addition to providing ω elements for feature evaluation statistics and manual external applications, CI The P system is internally via its feature function controller (Figures 2 and 3). Statistics can be used automatically. For example, the feature function controller is a partial phase Unnecessary by matching the relation and squared correlation values against a given cutoff value Features can be identified and excluded. Similarly, the feature function controller is a partial phase Identify redundant feature pairs by matching the squared difference between the values with a given cutoff value Can be When such an unnecessary or redundant feature is identified, the feature function The controller issues feature modification commands to the transducer input processor and It can be sent to the lance juicer output processor. Features In addition to modifying the transducer behavior during functional specification changes, the CIP system The system can also modify the elements of the coupling weight matrix ω. No need for ω element New, tuned joins with features, eg, feature f, one row down by row and column A weight matrix, such as ω {f, f}, can be created. line f and f The submatrix of ω that excludes columns is represented by ω <f, f>, and the deleted f row is ω <f>. Expressed by, an appropriate adjustment formula based on the standard matrix algebraic function is expressed by the following formula. ω {f, f} = ω <f, f> -ω <f><sup>T</sup>ω <f> / ω (f, f) (26) Parallel kernel operation As mentioned above for the CIP system in Figure 2, the parallel CIP kernel 21 Features process during parallel pass-through, monitoring and learning parameter memory update operation Use the bond weights between the sensors. As mentioned above, parallel CIP kernel 21 Processes F features per trial, along with the Mahalanobis Distance Processor 34 , F parallel feature processor 31<sub>l</sub>From 31<sub>F</sub>To use. The processor has a limited number of outputs that can be driven and is used to process feature values. Therefore, if many features are implemented, the number of outputs that the processor can drive is easily exceeded. I will get it. Parallel CIP systems access according to a coordinated timing scheme Features to be joy with switching junctions connecting processor pairs The output problem is solved by providing the access memory 23. Switchon In the junction, each processor exchanges related information and many processors are in parallel. To be able to operate. As will be described later, in CIP parallel kernel 21, the system The processor output only drives one output at any given adjustment time interval. is there. Conductor circuit layout and parallel car for F = 16 with reference to Figure 5. The wiring between the flannel processors is shown. According to the circuit shown, each feature processor is M. Connecting to the element of joint access memory 23, identified by the circle shown Also, depending on the circuit, each processor is every other through the systematic method described later. Paired with the processor of. Each processor 31<sub>1</sub>From 31<sub>16</sub>Also distance processor 3 Register D in 4<sub>1</sub>From D<sub>16</sub>It is connected to the. 16 features Processor 31<sub>1</sub>From 31<sub>16</sub>And besides the distance processor 34, Figure 5 shows the processor 31<sub>1</sub>From 31<sub>16</sub>And the conductor bus between the distance processor 34 Eout is illustrated. Figure 5 shows the processor 31<sub>1</sub>From 31<sub>16</sub>And join Wiring with access memory 23 and kernel control unit 32 is also shown. Has been done. The circuit shown in Fig. 5 is the lower bus layer, upper bus layer, lower and upper bus. A semiconductor layer that is in between and connects them, and a controller that is on top of all other layers. It can be implemented as a silicon chip layout that includes a trawl bus layer. Line 45L in Fig. 5<sub>1</sub>From 45L<sub>16</sub>Represents a bus in the lower bus layer, line 45 U<sub>2</sub>From 45U<sub>15</sub>Represents the bus in the upper bus layer, line E<sub>1</sub>From E<sub>16</sub>Join Lower layer bus and upper layer bus along the diagonal edge formed by the access memory element Represents an extension connection between, horizontal line 32<sub>1</sub>From 32<sub>29</sub>Is a joint access memo Kernel control unit to control switching of re23 Represents a joint access memory 23 control bus operated by 32. Continuing with reference to Fig. 5, each circle M (2,1), M (3,1) and M (3, 2) to M (16,15) are JAM memory and switching junctions, Represents a switching node that contains switching logic and memory registers. All of them can be arranged in the semiconductor layer between the lower bus layer and the upper bus layer. Features Processor 31<sub>1</sub>From 31<sub>16</sub>The circles M (16,1) to M (16,5) are connected Represents a coupled weight register in the processor that contains the main diagonal of the combined weight matrix and represents the distance. Circle D in detach processor 34<sub>1</sub>From D<sub>16</sub>Distance processor 34 and corresponding features Processor 31<sub>1</sub>From 31<sub>16</sub>Represents a register that communicates between. 16 lower buses 45L<sub>1</sub>From 45L<sub>16</sub>Corresponds to Features Processor 31<sub>1</sub>From 3 1<sub>16</sub>The distance processor 34 in the corresponding register D<sub>1</sub>From D<sub>16</sub>Connect to. Connection Extension E<sub>1</sub>From E<sub>16</sub>By each lower bus 45L<sub>1</sub>From 45L<sub>16</sub>Is its corresponding upper bus 45U<sub>1</sub>From 45U<sub>16</sub>Connected to. In addition, each lower bus 45L<sub>1</sub>From 45L<sub>16</sub>Corresponding features by processor 31<sub>1</sub>From 31<sub>16</sub>Is a dedicated joint access memory node M (2,1), M (3) as follows , 1) and M (3,2) are connected to M (16,16). Lower bus 45L<sub>1</sub>Is the bottom of M (16,1) at the bottom of JAM nodes M (2,1) and M (3,1) Connected via the part, lower bus 45L<sub>2</sub>Is from node M (3,2) to M (16,2) ) Connected to the bottom and extended to the upper bus 45U E<sub>2</sub>Node M (2) via , 1) is connected to the top. Similarly, the lower bus 45L<sub>3</sub>From 45L<sub>15</sub>Is each lower Connected to the bottom of the corresponding joint access memory 23 nodes and , Lower bus 45L<sub>1</sub>From 45L<sub>2</sub>And upper bus 45U<sub>2</sub>And E<sub>2</sub>About As such, the corresponding connection extension E<sub>3</sub>From E<sub>15</sub>Corresponding upper bus 45U via<sub>3</sub>From 45U<sub>15</sub>It is connected to the. Upper bus 45U<sub>3</sub>Are nodes M (3,2) and M ( It is connected to the top of 3,1). Each upper bus 45U<sub>2</sub>From 45U<sub>16</sub>Is the upper bus 45U<sub>3</sub>Along each upper bus as well It is connected to the top of each node. (Many connection extensions and nodes are illustrated All extensions E to make the drawing easier to understand<sub>3</sub>From E<sub>15</sub>Reference number is Not shown. If you are a person skilled in the art, the connection extension and the node should follow the rules used before. Can be identified using). Lower bus 45L<sub>16</sub>Is all the corresponding node M ( Connected to the top of M (16,15) from 16,1) and the corresponding extension E<sub>16</sub>To Corresponding upper bus 45U via<sub>16</sub>It is connected to the. Lower and upper buses and wiring are each memory carried out by CIP kernel 21 Contains a one-wire conductor for the bit. For example, the kernel is χ, ω and other If you use 32 bits to store the elements of a variable, each bus line will have its own bit Represents 32 conductors, one for each. Therefore, parallel kernel 21 is a feature processor. It is possible to communicate between the service and the access JAM element in parallel and quickly. Against this Then, each control bus 32 in Fig. 7<sub>1</sub>From 32<sub>30</sub>Represents three conductors The usage will be described later. Further refer to Fig. 6a and Fig. 6b for details of the switching junction. Show details, joint access memory node in kernel 21, joint access For components of memory node M (16,15) and M (16,15) bus connections This will be explained further. The components and connections considered here are the other components of the system. It also applies to board and bus connections. Figure 6a shows the flat of node M (16,15). A detailed surface view and associated switches and buses for the node are shown, with Figure 6b showing M (1). A side view of 6,15) and related switches and buses of the node are shown. 6th a As shown by the offset dashed line 805, the figure is slightly moved to the right for easy understanding. Fset lower bus 45L<sub>15</sub>And processor 15. This is shown in Fig. 5 There is no such offset, there the lower bus 45L<sub>15</sub>Is node M (16,15) It passes directly under the center of. Top and bottom as discussed in Figure 5 Club bus 45L<sub>16</sub>And 45L<sub>15</sub>Contains one conductor for each bit of memory accuracy There is. Similarly, the wires and switches described below have the same number of conductors and switches, respectively. Includes switch contacts. Figure 6a shows the details of the joint access memory node M (16,15). It includes the following: Memory cells containing ω (16,15), ω (16,15) ) Memory input switch S1, ω (16,15) to access memory output Power switch S2, and processor 31 at the output of S2<sub>16</sub>Upper bus 45L<sub>16</sub><sub></sub>And processor 31<sub>15</sub>Lower bus 45L<sub>15</sub>Dual switch S3 to join. No. Figure 8a shows the side view of the memory cell and the same three switches as in Figure 8a. Figure 6a is connected to the memory cell input including ω (16,15) via S2. Processor 31<sub>16</sub>Upper bus 45U<sub>16</sub>Is shown. Therefore, when S2 closes, ω ( Memory cells containing 16,15) are processors 31<sub>16</sub>Upper bus 45U<sub>16</sub>Including the contents of It will be updated so that it will be updated. Processor 31<sub>16</sub>Upper bus 45U<sub>16</sub>And processor 31<sub>15</sub>Lower bus 45U<sub>15</sub>Are interconnected when S3 closes and S2 opens. In addition, Rossessa 31<sub>16</sub>Upper bus 45U<sub>16</sub>And processor 31<sub>15</sub>Lower bus 45U<sub>15</sub>Is S Touch the output of a memory cell containing ω (16,15) when both 2 and S3 close Will be continued. Memory cells containing ω (16,15) when both S2 and S3 are closed The contents of are resident on both buses. The switches S1, S2 and S3 in Fig. 8a are control lines, respectively. Controlled by signals on C1, C2 and C3. These three controls The roll line includes the control bus line shown in Fig. 5. Any of these If the signal on the control line is positive, the corresponding switch closes. Similarly In addition, the switches S4 to S7 in Fig. 6a correspond to the control lines C4 to C. 7 Controlled by the signal above. Switches S4 and S5, respectively Processor 31<sub>15</sub>It is connected to the input bus 801 and the output bus 802 for The switches S6 and S7 are connected to the input bus 803 and the output bus 804, respectively. Has been done. Five basic switching operations ((a)-(e)) are performed and the joint operation is performed. The circuit shown in the cess memory 23 is realized as follows. (a) Processor 31<sub>15</sub>O And processor 31<sub>16</sub>Joint access to ω (16,15) by If S1 to S7 are open, closed, closed, closed, open, closed and open, respectively, (b) The variable value is sent from processor 16 to processor 15, in which case S1 to S7 Open, open, closed, closed, open, open and closed, respectively, (c) Processor 31<sub>15</sub>Ka Processor 31<sub>16</sub>Send the variable value to, in which case S1 to S7 are open, respectively. Open, closed, open, closed, closed and open, (d) ω (16,15) value processor 3 1<sub>16</sub>In that case S1 to S7 are open, closed, closed, open, open, closed and closed, respectively. Open and update the memory cell containing the (e) ω (16,15) value, in which case S 1 to S7 are closed, open, open, open, open, open and closed, respectively. Joint The timing of access memory switching will be described later. Parallel kernel 21 operations are coordinated by control unit 32 (a) Each feature processor 31<sub>1</sub>From 31<sub>16</sub>Is continuously busy during the trial (B) Each joint access memory bus has one variable value at any given time. (C) Each memory cell sends only one output value at any given time, (c) d) Each feature processor takes the output value of the processor as one input at any given time Will not be sent. Adjustment steps are sent to individual feature processors Control bus 32 with control signal<sub>1</sub>From 32<sub>29</sub>Controlled by Will be done. Continuing with reference to Fig. 5, in the calculation of χ when F = 16, from χ (1) Each χ (16) is a feature processor 31<sub>1</sub>From 31<sub>16</sub>Is calculated by. When calculating χ (16) from each χ (1), the sum between the product terms is calculated according to equation (13). It is calculated and the element along one line of ω is multiplied by the corresponding element of e [IN]. χ When calculating, the element of e [IN] has already been calculated (as described below). The elements of e [IN] are resident in feature processors 1 to 16, respectively. There are times when Each feature processor F first initializes the value of χ to 0 and then lowers the feature processor. And by accessing the joint access memory node along the upper bus Compute the χ (F) value of the feature processor. During each access, each feature processor Sasa F executes the following operation sequence. Third, the stored ω element is no Ingest along the do, and second, ingest the e [IN] element obtained at that node. , Third, multiply the two elements against each other to get the cross product, and fourth, the processor The loss product is added to the sum of movements of χ (F) performed by the processor. For example, focus Features Processor is Processor 31<sub>16</sub>And consider Fig. 5 Processor 31<sub>16</sub>Can access M (16,15). So of When processor 31<sub>16</sub>Multiplies ω (16,15) by e [IN] and moves its product χ ( Update the value of χ (16) it is calculating by adding to the value of 16) .. Also, processor 31<sub>15</sub>Accesses ω (16,15) and e [IN] (16 ), Multiply the values by each other and multiply the product by processor 31<sub>16</sub>The value of the movement χ (15) of Processor 31 by adding to<sub>15</sub>Updating the calculation of the value of χ (15) in To. FIG. 6 shows the control unit timing of the above-mentioned χ update step. The top signal shows the CIP system clock pulse as a function of time. Two graphs are switch controls along lines C1 to C7 as a function of time Indicates the value. At the time t between the first pulse and the second pulse, the switch is described above. Set according to switch operation (a), feature processor 31<sub>1</sub>From 31<sub>16</sub>To ω Send (16,15). At the time t + 1 between the second and third pulses, The switch is set according to the switch operation (b) described above, and the feature processor 31<sub>16</sub>Send e [IN] (15) to processor 31<sub>16</sub>Is ω (16,15) and e Add the product of [IN] (15) to the calculated movement value of χ (16). 3rd pulse and 3rd At the time t + 2 between the 4 pulses, the switch goes into the switch operation (c) described above. Therefore set and processor 31<sub>15</sub>Send e [IN] (16) to, then process Ssa 31<sub>15</sub>Is the product of ω (16,15) and e [IN] (16), which is the ohm of χ (15). Add to the calculated value. After the 4th clock pulse, switches S2 and S3 do it Open as shown by their corresponding C2 and C3 control values being low And processor 31<sub>15</sub>Bus and processor 31<sub>16</sub>Other without interference along the bus Update operation can be performed. Each processor calculates χ to the cross power Calculate the product, add the cross product to the moving χ sum of the processor, and each other processor The server is calculating another cross product and proceeds to add the product to the χ term of the processor. To do. Regarding the update of the element of ω, when updating ω (16,15) according to equation (11) In addition, ω [IN] (16,15), χ (15) and χ (16) are all one at the beginning. Obtained within the processor of. Next, one processor ω [O] according to Eq. (11) UT] (16,15) is calculated, after which the processor updates the updated value to ω (16,1). 5) Send to the storage cell. Control timing of operation update sequence with reference to Figure 8. Indicates The system clock pulse is illustrated as a function of time and the clock The four graphs below the pulse are the C2, C3, C6 and C7 lines as a function of time The control value along with is shown. Smell the time t between the first and second pulses The switch is set as described above for switch operation (d) and features Send χ (15) to Sessa 16. At the time t + 1 between the second and third pulses The switch is set and featured as described above with respect to switch operation (e). Send ω (16,15) to processor 16 and then processor 31<sub>16</sub>Is χ (15 ) And ω (16,15) are received, so the second term of Eq. (11) is calculated. Also processor 31<sub>16</sub>Calculates the value of χ (16) in advance and stores it internally. ( The operation of other processors that complete the 11 formula will be described later). number 3 After the clock pulse of, switches S2 and S3 have their corresponding C2 and And the C3 control value is freed to indicate low, processor 31<sub>15</sub>Bus and processor 31<sub>16</sub>Other update operations without interference along the bus It can be performed. Refer to Figure 9 to show how to tune χ and ω update processor operations. Su. Each entry in Figure 9 is pro at each interval in the overall χ or ω update process. Time that Sessa is running with itself or with one other feature processor Indicates the interval. Triangular table 911a is U<sub>1</sub>From U<sub>16</sub>Low and L<sub>1</sub>From L<sub>16</sub>column They have the rows and columns shown in the joint access memory in Figure 5. Correspond. As will be described later, the entry represents the control timing. Rote The Bull 911b is the feature of Fig. 5 Processor 31<sub>1</sub>From 31<sub>16</sub>The column corresponding to Have. As will be described later, the entry represents the control timing. Same number The entries in the table with are during the time interval indicated by the table entry. Represents a uniquely paired set of processors. Entry 911a is a χ update process Indicates which processor is paired during each interval. Features Paired at Arbitrary Time Intervals Processors are set in Figure 9. It is determined by searching for a time interval along Sabas or within its processor. For example, processor 31<sub>16</sub>Corresponds to the bottom row in port 911a in Figure 9. .. If you look at that row, during time interval 1, feature processor 16 and feature processor Sasa 2 puts ω (16,2) together with e [IN] (2) and e [IN] (16) It can be seen that it is used to update χ (16) and χ (2), respectively. this The control of the operation is the processor 31 described above with respect to Fig. 7.<sub>15</sub>And 31<sub>16</sub>Same as the control of. Similarly, feature processor 31<sub>16</sub>And special Signal processor 31<sub>3</sub>Updates χ (16) and χ (3) during interval 2 Feature processor 15 and feature processor 4 have χ (15) and χ during interval 2. (4) has been updated, and the same applies below. Entry 911b in Figure 9 shows which processor is performing the operation However, it is not paired with another processor during that particular time interval. For example, at time interval 1, processor 1 has e [IN] (1) xω [IN] (1). ) Is χ (1) processor 31<sub>1</sub>Update χ (1) by adding to the moving sum At the same time, processor 9 changes e [IN] (9) xω [IN] (9) to χ (9). Processor 31<sub>1</sub>Χ (9) is updated by adding to the moving sum. Professional Parallel CIP kernel by accessing Sessa and updating as described above Keeps all feature processors busy through the χ update processor Can be done. The numbers in Figure 9 can be used to identify the operating steps of the processor during the χ calculation. Form a systematic pattern. The following equation is kernel step 2 (i, f, g = 1,. .., F) During iteration i as part of the calculation of the matrix-vector product χ during the period Identifies processor g accessed by processor f. g (i) = [F-f + i] mod (F) +1 (27) For example, if F = 1, at time interval 2 (ie i = 2), the feature processor Ssa 31<sub>15</sub>(f = 15) interacts with processor 2 (g (i) = [ 16-15 + 2] mod (l6) +1 = [1] mod (16) +1 = 2). The number pattern in Figure 9 is for performing processor operations during the χ calculation. It also shows the systematic patterns of control lines that can be used. For example, 16 When using features, each given time interval number in the sequence is 911a in Figure 9. The value decreases along the line from the lower left boundary to the upper right boundary. As a result, three controls All corresponding lines along the line share the same timing. Line access memory nodes share the same set of three control lines can do. Therefore, the adjustment time interval pattern shown in Fig. 9 is shown in Fig. 5. It represents the control line pattern shown. The same adjustment as that expressed by Eq. (27) for calculating χ is the key to ω by the CIP system. Used to update the element, except for one. χ update shows the operation shown in Fig. 7. While the calculation of the operation is performed at the interval i of Eq. (27), the ω update is the operation shown in Fig. 8. The calculation of the ration is performed at the interval i in Eq. (27). After χ is calculated, χ (1) to χ (F) and e [IN] (1) to e [ The values of IN] F are resident in the feature processors 1 to F, respectively. next The Mahalanobis distance d is calculated according to equation (14) as follows. First, the product χ (1) xe [IN] (1) to χ (1) xe [IN] (F) from processor 1 Calculated simultaneously and in parallel by F, secondly, each such product is shown in Figure 7. Distance processor register D<sub>1</sub>From D<sub>F</sub>Sent to at the same time and in parallel, thirdly, the distance The rosessa according to Eq. (14), register D<sub>1</sub>From D<sub>F</sub>To find d by summing up the contents of Finally, the distance processor is in register D<sub>1</sub>From D<sub>F</sub>D to each feature processor via Back, the processor used it to update according to equations (10) to (12) Compute the variance and updated join weight matrix. Joint access memory and connected processor design are highly integrated circuits It is advantageous to carry out compactly. Comparable examples are each trial It is advantageous to have the optimum speed inside. Therefore, preferably kernel 21 As many features as possible on a single chip Realized by processors and JAM elements Revealed. Otherwise, the advantage of parallel processing speed is taken by serial communication between chips. It will be erased. Also, because the distance between internal components is short, between components High-speed electrical signal transmission can be obtained. The CIP parallel kernel also satisfies other design issues. (a) Signal deterioration problem (Widely known as fanout)-One feature processor of the JAM element Minimize the maximum number of inputs that can supply at any given time, and (b) space Utilization problem-Features Minimize the number of conductors required for communication between the processor and JAM elements Select. The CIP parallel kernel solves these various design problems as described above. Full through JAM bus and switching structure, along with column kernel feature processing adjustments Add. Those skilled in the art know that kernel 21 can be realized with analog circuits. U. An alternative analog embodiment is realized as follows. (a) Each analog JAM bus Is not a collection of digital bit wires, but a single conductor, (b) each digital The JAM switch has only one contact, (c) each digital JAM memory element It is not a large digital memory element but a small resistance-capacity network, and (d) arithmetic operations. A simple (non-sequential) analog circuit is used to perform. Furthermore, as mentioned above Some of JAM analog operations are paired with digital ALU operations Analog that is more compact, faster, and has acceptable accuracy by mating -You can create a digital hybrid. Sequential kernel operation The sequential kernel is the basic kernel operation described above for parallel kernels. Use yon. Therefore, whenever both kernels receive the same input, they will be sequential. Kernel operation produces the same output as parallel kernel operation Is done. However, sequential operation is only one, not F processors It is generally slow because it is obtained using the processor of. Some sequential power Kernel operations simulate parallel kernel operations in exactly the same way. It is done differently for efficiency, not for rate. Χ calculation step and ω [OUT] calculation step of sequential kernel operation The operation is performed for optional storage devices and speeds. Both steps ω ( Including 1,1), followed by ω (2,1), followed by ω (2,2), followed by ω (3,1) Then, ω (3,2) continues, ω (3,3) continues, and then ω (F, F) continues. It is based on storing the element of ω as a symbol string. χ calculation step and The ω [OUT] calculation step goes from the first element to the continuously stored elements of ω. Access to later elements. The overall effect is a nested loop with both steps like this Is to be much faster than if it were done in the traditional way using. χ And for a sequence of operations that calculate the ω update step, it Consider each in relation to Figures 19 and 20, respectively. Parallel system operation Ingress Transducer Operation 504, Kernel, See Figure 10. Calculation 505, learning weight controller operation 508, feature function controller La operation 511, transducer output operation 506 and Raffic Display Operation 517 separate subsystems, system It can be executed at the same time at the level. There are two types of operations It is a parallel operation and a management operation. Concurrency operation Young is shown in 503, 504, 505, 506, and 507 and is occasionally managed. Perations are shown in 502,508,511, and 517. Parallel, running quickly during each trial, with reference to Figures 10 and 2. The operation is performed by the transducer input processor 24. Executed by suser input operation 504, kernel processor 24 Kernel operation 505 and transducer output processor 26 Includes Transducer Output Operation 506 to be performed more. Back The storage device uses the output value from each device as the input value to the next device. Allow both devices to process data for different trials at the same time Will be done. Transducer input operator by using buffering Shon 504 creates input features for the next trial and transformer Measures passed on from the trial preceded by Jucer Output Operation 506 Kernel operation 505 is for parallel trials It is possible to create a kernel output function and a learning parameter update function. Management operations that can be performed occasionally over a period of several trials Learning weight control operation performed by the learning weight controller 40 508, feature function controller 41 executed by feature function controller 41 Peration 511 and Graphic Display Operation 517 Includes. Buffering enables parallel management operations and Parallel concurrency operations performed by coordinator 38 by buffering It becomes possible. During the management operation, the output information from kernel 21 is Inn 40<sub>1</sub>From 40<sub>F</sub>With input information to the learning weight controller 40 via, 40d And the output information from kernel 21 is line 41<sub>1</sub>From 41<sub>F</sub>The feature function controller can be used via. Buffers are parallel Management operation based on previous trial statistics while operation continues Used to allow the action to proceed in the buffer. The value monitored during operation 510 within the learning weight controller 40 is learned Sent to weight control operation 509. Features Function controller 41 Within, the output from the kernel is processed during feature function operation 513 and operated. It is sent to the control feature function during session 512. After the kernel 21 output value is received in the buffer, the feature function controller 41 Can execute equations (23) to (25) and from equations (17) to (19) You can execute expressions. Features Function Controller Operation 508 and During learning weight controller operation 511, kernel 21 operates its parallel operations Continue the ration. Therefore, three operations 504,505, And 506 proceed simultaneously. CIP systems are sometimes featured, along with parallel pass-through, monitoring and learning operations Monitor functionality. The learned join weight value and the learned amount are included in the monitoring operation. The scattered value is received from kernel 21, the feature multiple correlation value is calculated according to Eq. (23), and ( Calculate the feature tolerance band ratio value according to Eq. 24), and calculate the partial correlation value according to Eq. (25). Includes calculation. After calculating the feature function monitoring statistics, the statistics (a) Manuscript to CIP users via graphic display operation 517 Provides al-interpretation and refinement, or (b) automatic refinement of CIP statistics 51 2 Can be used for operations. Monitoring feature Function Operation 513 result correction Switching operation 5 The feature specifications are changed by controlling 07. Switch operation 5 By triggering 07, the CIP system 10 operates as described above. Reinitialize the feature specifications via 502. As mentioned above, the CIP system takes advantage of the quantity count easy Bayes. In order to use, the feature function monitoring statistics that satisfy Eqs. (23) to (25) have been learned. It can be interpreted directly, along with the mean and the learned join weights. trial Each learning weight from 1 to trial t, l (1) to l (t), is the previous ton as follows: It shall be the ratio of the amount count of each trial to the amount count of the trial. Then, the quantity count for the initial parameter value is a feature from trial 1 to t. Both are displayed as q (0) to q (t), and the following equation is obtained.<img file="JP2000511307A_D0019.tif" />Furthermore, the initial mean vector is the mean of the q (0) quantity initial vector, and the tria Input feature vectors m [IN] from le 1 to t are set to trial 1, respectively. The average of the q (t) quantity vector from the q (1) quantity vector to the trial t. It shall be a value. Then, all regression parameter values learned in parallel from statistical theory And all refinement parameters that can be used in parallel are q (0) + q (1) + ... Based on an equally weighted quantity count from the total sample size of + q (t) It can be interpreted as an arithmetic mean statistic. For example, the learned feature mean vector at the end of trial 10 is q (0) + q (1) Has an interpretation of the mean of the quantity values of + ... + q (10), and any other trial The learned feature mean vector after the number has the same interpretation. As a second example, q (1) = q (2) = q (3) = R, where R is a positive constant Impact sequence that satisfies equations (20) to (22) by setting, N In that case, from algebra based on equation (28) to equation (20) From l (1) = 1 / R, l (2) = 1 / (R + 1), l (3) as shown in equation (23) ) = 1 / (R + 2) etc., and R = q (0) / R<sub>1</sub>Is. Therefore, (28 ) From equation (20) to equation (23), etc. Importance CIP sequence deviation Not only that, the quantitative observational interpretation of equal importance learning by the CIP system can be obtained. Uses the quantity count Easy Bayes used by the CIP system As a result, the feature regression parameters and feature function monitoring parameters are set for each trial. Traditional statistical procedures and traditional neurocomputing that can be evaluated in parallel It can be carried out more easily than the parameters obtained by the procedure. As mentioned above, parallel CIP systems use buffered communication to be faster. It can operate quickly. Also, as mentioned above, as many CIPs as possible Significant inter-chip communication time loss by performing peration on one chip Can be avoided. As a result, some CIP subsystems are single Conducted on various layers of the chip and communicated between layers by parallel buffering By doing so, the overall CIP operating speed can be maximized. Figure 11 shows the characteristics of the parallel kernel accessed by the learning weight controller 40. Processor 31<sub>1</sub>From 31<sub>F</sub>Memory position in 311, learning weight controller buff 1101, and the corresponding bus to the learning weight controller buffer 1101 1<sub>1a</sub>From 41<sub>Fd</sub>Is shown. As shown in FIG. 11, the buffer element 111 is a kernel. 21 Geometric configuration similar to memory location, corresponding to each parallel kernel memory location It can be aligned in parallel with the kernel memory location as shown. Buffer power The buffer is a parallel car with the minimum wiring due to the parallel structure for the memory 21 memory position. It can be provided in the upper or lower layer of flannel. The buffering to the feature function controller is shown with reference to Fig. 12. 1st 2 Figure shows the parallel kernel feature pro accessed by feature function controller 41. Memory location 411 in the sesser, feature function controller buffer 1201, and Features Function Corresponding bus to controller buffer 1201 41<sub>JAM (2,1)</sub>From 41<sub>JAM (F, F)</sub>And 41<sub>1</sub>From 41<sub>F</sub>Is illustrated. As shown in Fig. 11, the first The buffer elements in Fig. 2 are configured in the same way as the kernel 21 memory location, and each parallel device It corresponds to the kernel memory position and is aligned in parallel with the kernel 21 memory position. Ba The buffer can be placed in the upper or lower layer of the parallel kernel with minimal wiring requirements. To. Geometrically aligned buffering and multi-layer chips in Figures 11 and 12 Combined with the design, the CIP subsystem is single-chip or aligned. Can reside on a single array of apps. As a result, waste in the system Communication time is minimized. Sequential system operation A sequential kernel subsystem is shown with reference to Figure 13. Sub in Fig. 13 Vectors and parameters transferred between systems are illustrated in more detail than in Figure 3. It has been. Learning weight controller 40a, feature function controller 41a and -The various inputs and outputs of the dinator 38a are illustrated. Be monitored Various parameters and control functions performed by the sequential system are related It is the same as a parallel system except that the data to be transferred is sequentially performed. One CIP operation at a time using one available processor Because it is not executed, the CIP system is all related to the parallel system mentioned above. Operations can be performed, though not very fast. Also, mosquitoes Subsystem because only one processor can be used for channel operation Simultaneous operations like the parallel kernel example are not performed at the level. However, except for speed issues, the CIP system is on par with sequential implementation. It is as powerful as it would be if implemented using a column processor. Also sequentially CIP The examples have at least two advantages over the parallel examples, and the sequential examples are special. Generally costly because it can be implemented within a traditional computer rather than a parallel circuit of design Less and more than parallel computers supply in specialized circuits The features per trial can be supplied by a conventional computer. That As a result, in the sequential CIP example, the trial generation rate is the number of features per trial. It is useful in relatively low applications. Alternative kernel example Alternative operations for the kernel include (1) David-Fetcher-Pow. Updated coefficient matrix used by L (DFP) numerical optimization algorithm , (2) Multiply the symmetric matrix by the vector, (3) Feature function control Adjust the combined weight matrix of features removed in, and (4) kernel It involves learning and becoming an input transducer. 4 responses related to: Consider everything for based on kernel modifications. Starting with a numerical optimization application, the DFP method is the maximum for some variable functions ( Or one of several iterative ways to find the (or minimal) independent variable. number Value optimization methods are generally useful but generally slow. For example, numerical optimization The conversion method is used to find the optimal value associated with the 5-day weather forecast, but super Even computers generally take hours to converge. Numerical optimum value method In, the DFP method learns derived information during an iterative search process that is not readily available. Therefore, it is particularly useful in various applications. Parallel kernel processes are used to implement high-speed parallel information processing systems As its modified version can be used for new fast numerical optimum systems it can. In particular, for sequential DFP updates based on F independent variables to converge to the optimal solution If it takes s seconds, parallel DFP updates only take about s / F seconds to converge .. For example, the 5-day weather forecast needs to optimize the function of 2,000 variables, It shall take 20 hours to converge using the conventional (sequential) DFP method. The same optimization problem can be solved by parallel processing corresponding to the DFP method similar to the parallel kernel. If it comes, it will converge in about 18 seconds. In the DFP method, the inverse element of the matrix is continuously updated as part of the normal operation. Instead of updating the inverse element of the covariant matrix as in the CIP system, the DFP system Algorithm is the inverse of the estimated matrix of quadratic derivatives called the information matrix. Update the original. The formula for updating the DFP inverse element is different from the formula for updating the CIP inverse element However, updating the DFP with an extension of the parallel CIP kernel algorithm Can be done. The DFP information matrix inverse element update formula is ω [DFP, OUT] = ω [DFP, IN] -c [DFP] χ [DFP]<sup>T</sup><sup></sup> χ [DFP] + b [DFP] y [DFP]<sup>T</sup>y [DFP ] (29) And here, c [DFP] = 1 / d [DFP] (30) χ [DFP] = e [DFP] ω [DFP, IN] (31) d [DFP] = e [DFP] ω [DFP, IN] e [DFP]<sup>T</sup> (32) b [DFP] = χ [DFP] y [DFP]<sup>T</sup> (33) DFP update formulas (29) to (33) should be adapted to suit DFP updates Can be done. In particular, the DFP corresponding to the kernel process is the parallel CIP kernel 21. So that the parallel CIP kernel computes equation (11) using the same number of steps as Calculate equation (29) and calculate the distance function that the parallel CIP kernel satisfies equation (14) The distance function that satisfies Eq. (32) and its inner product that satisfies Eq. (33). The matrix on which the function is calculated and the parallel CIP kernel 21 satisfies equation (13) Sum the matrix-vector product that satisfies equation (31) to calculate the cross product Calculate. The difference between the two parallel methods is the DFP constants in Eqs. (A) and (29). It is simpler than the parallel CIP kernel expression (12) and supports kernel processes. DFP solves one parallel CIP kernel product to compute equation (14) Solve the two inner products to calculate equations (32) and (33) instead of 11) Sum one corresponding parallel CIP kernel cross product to calculate the second term of equation Instead of calculating, the second term of Eq. (29) and two cross product to compute the beauty paragraph 3 Is to calculate the term of. More computational complexity than vector multiplication of symmetric matrix is repeated and fast An example of an adapted kernel with less is obtained. The kernel example is (13). The formula is an example, the calculation of the product saves only the operations necessary for the calculation and excludes everything else. It can be simplified by doing. Parallel CIP kernel examples and other With all adapted versions, using parallel processing instead of sequential processing Results can be obtained at F times faster speed. For the adapted corresponding kernel in the CIP system, equations (a) and (26) The constant coefficient of the second term of the above does not use the distance function, and is the most as shown in equations (b) and (13). Initially without the need for a matrix-vector product, just the outer product of two vectors (2) 6) The feature exclusion adjustment formula (26) is a kernel update in that the second term of the formula is calculated. A simplified version of equation (11). As a result, the parallel CIP kernel Eq. (26) can be solved by simplification. Adapted parallel CIP carne for feature modification and input transducer processing Student Input Transducer is the first useful feature when it comes to Le Operation "Teached" to use only, then operate to create features Can be used. For example, the CIP system has features 1 to 100. One independent variable value feature 1 needs to be predicted as a function of the independent variable value And. Feature 2 by modifying the kernel process during a traditional series of learning trials To identify 99 optimal join weights that pass on feature 1 from to feature 100 You can learn. Corresponds to feature 2 to feature 100 after learning Of an input transducer with 99 inputs and the only output corresponding to feature 1. You can use the learned module instead. With an input transducer When used as a module, its learning and updating operations are bypassed. It differs from the kernel in that it is struck. Therefore, such a module The only kernel modification required to implement is learning during feature pass-through operations Input for learning vs. feature pass-through operation 2 with a small amount of internal logic to bypass It is a forward indicator. CIP system process The parallel CIP kernel system implemented in the present invention with reference to FIGS. 14 to 17. The preferred steps of the stem are shown. Process steps shown in Figures 14 to 17 All uploads are done within the parallel kernel 21 subsystem of CIP system 10. In step 1400, as mentioned above, kernel 21 initiates initial processing. When the learning weight l, the feature vector m [IN] and the viability vector ν are set. Receive. In step 1401, the kernel has an average vector μ [IN], ω Learning parameters containing the values of [IN], ν [D, IN], λ [IN], and μ [IN] Access the meter memory element.
[IN], ω [IN], ν [IN], and Each value of 1 [IN] is stored in the training parameter memory during the previous trial. Calculated as force. However, whether this is an initial CIP system iteration For example, the μ [IN] value is zero, the ω [IN] value corresponds to the identity matrix, and ν [ Not only the D, IN] value but also the μ [IN] value is zero. In step 1402, the kernel is viable according to equation (3). Component features from tor learning weight, global learning weight 1 and learning history Calculate the parameter l [IN] value. The process then proceeds to step 1404. In step 1404, the feature average vector m [IN] is changed according to Eq. (2). Will be renewed. In step 1406, the intermediate pass-through feature vector according to equation (9) e [IN] is updated. In step 1408, study shoes according to equation (4) The history parameter l [OUT] is updated. The process then proceeds to B in Figure 15. Preferred Steps of the Process of Preferred Examples of the Invention with reference to FIG. Continue to consider. In Fig. 15, the intermediate matrix / vector product is calculated as described above. The preferred process to be performed is shown. In step 1500, intermediate matrix / Each element χ (f) of the vector product is initialized to zero and the process goes to step 1502 move on. In step 1502, the kernel has the adjustment time described above with respect to FIG. Start accessing the processor pair according to the method. In step 1504, each Processor vs. f is the proper coupling weight at the joint access switching node Search only ω [IN] (f, g). Similarly, as described above with respect to FIG. The appropriate intermediate pass-through feature value e [IN] (g) is searched for in step 1506. .. As mentioned above, in step 1508, the element χ (f) depends on the cross product. Is incremented. The process proceeds to step 1510, where the final adjustment time interval of the adjustment time method It will be confirmed if it has been reached. If the final adjustment time interval has not been reached, the process In step 1512 proceed to the next time interval and then from step 1502 P1508 is repeated again. By repeating steps 1502 to 1508 χ The moving sum of the calculation in (f) is obtained. During the last connection time in step 1510 If the distance is reached, the calculated value is the distance pro in step 1552 as described above. Stored in Sessa. The process then proceeds to C in Figure 16. In the present invention, the output value of the kernel subsystem 21 is calculated with reference to FIG. The steps of a preferred embodiment are shown. In step 1602, according to equation (17) The output regression feature vector m [IN] is calculated. Next, in step 1604, the pass-through feature vector m [OU] according to Eq. (16). T] is calculated. The process then proceeds to step 1606. Step 1606 In the distance processor at the distance ALU34, according to Eq. (8), Maharano Calculate the screw distance value. As mentioned above, each processor depends on a particular processor The calculated distance value cross product is stored in the distance processor. All processors After receiving each cross product value from, the distance processor is given by the feature processor The distance measure d is obtained by summing all the distance values obtained. Then the process is D in Figure 17. Proceed to. With reference to FIG. 17, a binding weight matrix is required in a preferred embodiment of the present invention. The steps of the process of updating the element ω (f, g) are shown. Step 1702 smell Then, the variance ν [D] (f) is calculated according to Eq. (10). Then the process is Proceed to pp 1706, where pro according to the coupling time scheme described above for FIG. Processor g is accessed by Sessa F. In step 1710 It is checked whether the Rossessa g is accessed by the lower bus line. ( See Figure 8). The processor is accessing processor g via the lower bus If so, the process of processor f proceeds to step 1712. Step 17 At 12, the intermediate matrix / vector product χ (g) is checked against processor f. Be searched. In step 1714, the node of the currently paired processor The appropriate coupling weight element corresponding to the memory element located at is pro according to Eq. (11). Updated by Sessa f. At step 1720, the kernel reaches the final connection time interval in the tuning time scheme. Check if you did. If the final interval of the adjustment time method has not been reached, At 1722, the kernel advances to the next adjustment time interval. Continue to step 1722 Then steps 1706 to 1720 are repeated. In step 1710, processor g is accessed via its upper bus. If so, the processor f will generate a parallel trial output and enter for the next trial. Read the force. The process goes from step 1716 to step 1720 to it The above was mentioned. In step 1720, the adjustment time method last connection time If the gap is reached, kernel functionality will end for the current trial. Refer to FIGS. 18 to 21 for the process performed in the present invention. Here are the preferred steps for the sequential CIP kernel 21a. Figures 18 to 21 All process steps shown in are subsequential kernel 21a of CIP system 11 It is done in the system. As mentioned above, in step 1800, the sequential carne Le 21a has a learning weight l, a feature vector m [IN] and a feature vector m [IN] at the start of initial processing. Receives the viability value ν. In step 1801, the kernel is μ [I Learning parameters including N] value, ω [IN] value, ν [D, IN] value and λ [IN] Access the data memory element.
[IN], ω [IN], and l [IN] The value is calculated as output stored in the training parameter memory during the previous trial ing. However, if this is an initial CIP system iteration, then each μ [IN] (f) The value is equal to zero, the ω [IN] value corresponds to the identity matrix, ν [D, I Not only the N] value but also the λ [IN] value is 1. In step 1802, the kernel is viable according to equation (3). Component characterization from tor ν, global learning weight l, and λ [IN] values Calculate the training weight. In step 1804, the feature average vector according to equation (2). Torr is updated. In step 1806, the intermediate pass-through feature according to Eq. (9) The vector e [IN] is calculated. In step 1808, according to equation (4) The learning history parameter λ [OUT] is updated. In the present invention to calculate the intermediate matrix / vector χ with reference to FIG. Here are the preferred steps of the sequential CIP kernel process performed in. 1st 9 Traditional double loop (ie, matrix) by the process of considering the figure Loop for all low values in and loop for all column values in the matrix A method of calculating the intermediate matrix / vector χ without executing p) is provided. Ma Trix ω element goes from ω (1,1) to ω (2,1) to ω (2,2) ω (3,1) To ω (3,2) to ω (3,3) and so on, in a continuous order corresponding to ω (F, F) It is stored as a symbol string of. In step 1902, the position h of the first ω element Is initialized to 1. In step 1904, intermediate matrix / vector χ Is set to zero. The process then proceeds to step 1906, where the matrix The low value f corresponding to the ω element stored in the format is set to zero. Then the process Goes to step 1912, where the low value is incremented by 1. Step 1908 In, the column value g corresponding to the ω element stored in the matrix format is then zero. Is set to. Then the process goes to step 1914 where the column value g is 1. Is incremented. The process then proceeds to step 1916, where the intermediate matrix / Column g is the moving sum that calculates the product of the vector product ω (f, g) and the corresponding χ (f). Is incremented by the current intermediate pass-through feature vector e [IN] (g) to. In step 1920, check if the column value g is less than the low value f Will be done. The corresponding column value g is not less than the low value f and is an ω (f, g) element If indicates that is on the main diagonal of the bond weight matrix, the process steps Proceed to 1924, where the element position h is incremented by l. Then the process is stepping Proceed to 1930, where the low value f is equal to the column value g and the ω element is the bond weight matrix. It is confirmed whether it indicates that it is on the main diagonal of the kusu ω [IN]. here And the column value is smaller than the low value, and more elements of ω are included in the corresponding row If it indicates that it is, the process is step 1914, as described above. Proceed to where the column value g is incremented by 1 and the process proceeds to step 1916. .. In step 1920, if the low value is equal to the column value, then step 192 In 2, the moving sum for calculating the intermediate matrix / vector product χ (g) is the current ω Incremented by the amount obtained by multiplying the (f, g) element by the intermediate pass-through feature vector e (f) of column f Will be done. The process then proceeds to step 1924, where it accesses the next ω element. Therefore, the position h of the ω element is incremented by 1. Low value in step 1930 If is equal to the column value, the process proceeds to step 1940. Step 194 At 0, it is checked if the low value is equal to the total number of system features. Low If the values are not equal to the total number of features, indicating that not all ω values have been evaluated The process proceeds to step 1912 and continues as described above. Step 194 At 0, the b value is not equal to the number of features, and all ω values are evaluated. If indicated, the process proceeds to F in Figure 20. A book that calculates the output value of the sequential kernel subsystem 21a with reference to Fig. 20. The steps of a preferred embodiment of the invention are shown. In step 2002, (17) The output regression feature vector is calculated according to the equation. Next, in step 2004, the pass-through feature evaluation m [OUT] is calculated. Next The process proceeds to step 2006. Distance ALU in step 2006 Calculates the Mahalanobis distance in the distance processor according to equation (8). Said As you can see, each processor uses the distance calculated by a particular processor as a distance processor. Store in. Distance processor after receiving each cross product from all processors The CPU sums all the distance values given by the feature processor to obtain the distance measure d. Next The process proceeds to G in Fig. 21. Refer to Fig. 21 and step the ω update process for the sequential kernel process. Show. In step 2102, the position of the ω element in the sequence of the ω element The setting is initialized to 1. Corresponds to the ω element in the symbol string in step 2104 The low value is initialized to zero. Then the process goes to step 2106, where The low value f is incremented by 1. Then the process goes to step 2108, where the minutes The scattering element ν [D, OUT] is updated with respect to the current low f-number. Then the process is Proceed to step 2112 where the column value g is initialized to zero. Step 21 At 14, the column value g is incremented by 1. Then the process is step 211 Proceed to 6, where the coupling weight element of ω is updated according to equation (11). Next, the process Su goes to step 2118. In step 2118, the next ω required in the symbol string The ω position value h is incremented by 1 to access the element. At step 2120 It is checked whether the column value g is equal to the low value f. Column value g is low value If it is not equal to f, the ω element corresponding to that row will not be updated any further, as described above. As such, the process proceeds to step 2114 and steps 2116 and 2118. To be completed. If the column value g is equal to the low value f in step 2120, then the main diagonal The process proceeds to step 2130, indicating that the ω element corresponding to is reached. In step 2130, it is checked whether the low value f is equal to the number of features. B -If the values are not equal to the number of features, the process proceeds to step 2106, where it is low. The value is incremented and the process proceeds as described above for the previous step. Step If the low value f is not equal to the number of features at 2130, the update process is terminated. The kernel function for the current trial is terminated. FIG. 22 shows preferred embodiments of the present invention for system monitoring. In step 2202, the learning feature coupling weight ω [OUT] and the learning feature variance ν [D, OUT] is received from the kernel. Then the process is step 2203 Go to and calculate the feature multiple correlation c [M] according to Eq. (23). To step 2204 Then, the tolerance band ratio r is calculated according to Eq. (24). In step 2206 Then, the partial correlation is calculated according to Eq. (25). In addition, the CIP system is on To monitor the force deviation, anomalous deviations are detected in the system input as described above. And learning can be enabled in step 2208. Step 22 At 10, the Mahalanobis distance can be drawn on the output display monitor 14. Calculate and display the standard deviation measure according to equations (18) and (19) can do. Clean any desired output of the CIP system according to user specifications It can be displayed on the Tep 2212 and evaluated by the user. Although preferred embodiments of the present invention have been described, the issues specified in the claims have been described. Various changes can be made within the Ming dynasty.
12 members in 8 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 08333204 | United States of America | – | |
| 33320494 | United States of America | A | |
| 33320494 | United States of America | A | |
| 9514160 | United States of America | W | |
| 9514160 | United States of America | W | |
| 333204 | – | – | – |
| PCTUS199514160 | – | – | – |
| US19940333204 | – | – | – |
| WO1995US14160 | – | – | – |
Members12
| Document | Office | Kind | |
|---|---|---|---|
| CA2203776A1 | Canada | A1 | |
| WO9614616A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU4141396A | Australia | A | |
| EP0789869A1 | European Patent Office (EPO) | A1 | |
| US5835902A | United States of America | A | |
| AU706727B2 | Australia | B2 | |
| CN1229482A | China | A | |
| NZ296732A | New Zealand | A | |
| JP2000511307AThis record | Japan | A | |
| US6289330B1 | United States of America | B1 | |
| CA2203776C | Canada | C | |
| EP0789869A4 | European Patent Office (EPO) | A4 |
Numbers
- Publication
- 2000-511307
- Publication, DOCDB
- 2000511307
- Publication, EPODOC
- JP2000511307
- Application
- 8515395
- Application, DOCDB
- 51539596
- Application, EPODOC
- JP19960515395
Titles2
- Japanese
- 【発明の名称】同時学習およびパフォーマンス情報処理システム
- English
- [Title of Invention] Simultaneous learning and performance information processing system
Classification
- CPC, 4
- G06N3/063
- G06N3/10
- G06V10/771
- G06F18/2113
- IPC, 6
- G06N99 00
- G06F15 18
- G06N3 063
- G06N3 08
- G06N3 10
- G06V10 771