Process for exchanging information in a multiprocessor system.
30 claims: 14 independent, 16 dependent
- 1Système multiprocesseur, du type comprenant une mémoire centrale (RAM) organisée en blocs d'informations (bi), des processeurs de traitement (CPU₁... CPU j ... CPU n ), une mémoire-cache (MC j ) reliée à chaque processeur de traitement (CPU j ) et organisée en blocs d'informations (bi) de même taille que ceux de la mémoire centrale, un répertoire (RG j ) et son processeur de gestion (PG j ) associé à chaque mémoire-cache (MC j ), ledit système multiprocesseur étant caractérisé en ce qu'il est doté :. d'un ensemble de registres à décalage, dit registres-mémoire (RDM₁... RDM j ... RDM n ), chaque registre (RDM j ) de cet ensemble étant associé à une logique de transfert (TFR j ) et connecté à la mémoire centrale (RAM) de façon à permettre, en un cycle de cette mémoire, un transfert parallèle en lecture ou écriture d'un bloc d'informations (bi) entre ledit registre et ladite mémoire centrale, . des registres à décalage, dits registres-processeur (RDP₁... RDP j ... RDP n ), chaque registre à décalage processeur (RDP j ) étant associé à une logique de transfert (TFR' j ) et relié à la mémoire-cache (MC j ) d'un processeur (CPU j ) de façon à permettre un transfert parallèle en lecture ou écriture d'un bloc d'informations (bi) entre ledit registre à décalage (RDP j ) et ladite mémoire-cache (MC j ), . un ensemble de liaisons séries (LS₁... LS j ... LS n ), chacune reliant un registre à décalage mémoire (RDM j ) et un registre à décalage processeur (RDP j ) et adaptée pour permettre le transfert de blocs d'informations (bi) entre les deux registres considérés (RDM j , RDP j ) à une fréquence de transfert F au moins égale à 100 mégahertz, . des moyens de communication d'adresses de blocs comprenant un registre à décalage complémentaire (RDC j ) connecté sur chaque liaison série (LS j ) en parallèle avec le registre à décalage mémoire correspondant (RDM j ) de façon à permettre la transmission des adresses par les liaisons séries et leur chargement dans lesdits registres à décalage complémentaires (RDC j ), un arbitre de gestion d'accès (ABM) étant relié auxdits registres à décalage complémentaires (RDC j ) et à la mémoire centrale (RAM) en vue de prélever les adresses contenues dans lesdits registres (RDC j ) et de gérer les conflits d'accès à la mémoire centrale (RAM), la logique de transfert (TFR j ) associée à chaque registre-processeur (RDP j ) étant contrôlée par le processeur de gestion (PG j ) de façon à permettre la transmission des adresses via ledit registre-processeur (RDP j ).
- 2Système multiprocesseur, du type comprenant une mémoire centrale (RAM) organisée en blocs d'informations (bi), des processeurs de traitement (CPU₁... CPU j ... CPU n ), une mémoire-cache (MC j ) reliée à chaque processeur de traitement (CPU j ) et organisée en blocs d'informations (bi) de même taille que ceux de la mémoire centrale, un répertoire (RG j ) et son processeur de gestion (PG j ) associé à chaque mémoire-cache (MC j ), ledit système multiprocesseur étant caractérisé en ce qu'il est doté :. d'un ensemble de registres à décalage, dit registres-mémoire (RDM₁... RDM j ... RDM n ), chaque registre (RDM j ) de cet ensemble étant associé à une logique de transfert (TFR j ) et connecté à la mémoire centrale (RAM) de façon à permettre, en un cycle de cette mémoire, un transfert parallèle en lecture ou écriture d'un bloc d'informations (bi) entre ledit registre et ladite mémoire centrale, . des registres à décalage, dits registresprocesseur (RDP₁... RDP j ... RDP n ), chaque registre à décalage processeur (RDP j ) étant associé à une logique de transfert (TFR' j ) et relié à la mémoire-cache (MC j ) d'un processeur (CPU j ) de façon à permettre un transfert parallèle en lecture ou écriture d'un bloc d'informations (bi) entre ledit registre à décalage (RDP j ) et ladite mémoire-cache (MC j ), . un ensemble de liaisons séries (LS₁... LS j ... LS n ), chacune reliant un registre à décalage mémoire (RDM j ) et un registre à décalage processeur (RDP j ) et adaptée pour permettre le transfert de blocs d'informations (bi) entre les deux registres considérés (RDM j, RDP j ) à une fréquence de transfert F au moins égale à 100 mégahertz, . des moyens de communication d'adresses de blocs comprenant un bus commun de communication parallèle d'adresses de blocs (BUSA) reliant les processeurs (CPU j ) et la mémoire centrale (RAM) et un arbitre de bus (AB) adapté pour gérer les conflits d'accès audit bus, le processeur de gestion (PG j ) étant relié au bus commun (BUSA) de façon à permettre la transmission des adresses sur ledit bus commun.
- 3Système multiprocesseur selon l'une des revendications 1 ou 2, caractérisé en ce que :- chaque registre à décalage mémoire (RDM j ) et chaque registre à décalage processeur (RDP j ) sont dédoublés en deux registres, l'un spécialisé pour le transfert dans un sens l'autre pour le transfert dans l'autre sens, - chaque liaison série (LS j ) comprend deux liens séries unidirectionnels de transfert bit à bit, reliant le registre à décalage mémoire (RDM j ) dédoublé et le registre à décalage processeur correspondant (RDP j ) dédoublé, ces liens étant connectés auxdits registres pour permettre, l'un un transfert dans un sens, l'autre un transfert dans l'autre sens.
- 4Système multiprocesseur selon l'une des revendications 1 ou 2, caractérisé en ce que chaque liaison série (LS j ) comprend un lien bidirectionnel de transfert bit à bit, connecté au registre à décalage mémoire (RDM j ) et au registre à décalage processeur correspondant (RDP j ) et une logique (LV) de validation du sens de transfert de façon à permettre un transfert alterné dans les deux sens.
- 5Système multiprocesseur selon l'une des revendications 1, 2, 3 ou 4, comprenant des moyens de gestion des données partagées entre processeurs, en vue d'en assurer la cohérence.
- 6Système multiprocesseur selon la revendication 5, caractérisé en ce que les moyens de gestion des données partagées comprennent :. un bus spécial de communication parallèle de mots (BUSD) reliant les processeurs (CPU j ) et la mémoire centrale (RAM), . une logique de partition (LP j ) associée à chaque processeur (CPU j ) et adaptée pour différencier les adresses des données partagées et celles des données non partagées de façon à transmettre celles-ci sur les moyens de communication d'adresses avec leur identification, . une logique de décodage (DEC) associée à la mémoire centrale (RAM) et adaptée pour recevoir les adresses avec leur identification et aiguiller les données en sortie mémoire soit vers le registre à décalage mémoire correspondant (RDM j ) pour les données non partagées, soit vers le bus spécial de communication de mots (BUSD) pour les données partagées.
- 7Système multiprocesseur selon la revendication 5, caractérisé en ce que les moyens de gestion des données partagées comprennent, d'une part, un bus spécial de communication parallèle de mots (BUSD) et un bus spécial commun de communication d'adresses de mots (BUSAM) reliant les processeurs (CPU j ) et la mémoire centrale (RAM), d'autre part, une logique de partition (LP j ) , associée à chaque processeur (CPU j ) et adaptée pour différencier les adresses des données partagées et celles des données non partagées, de façon à aiguiller les premières vers le bus spécial commun (BUSAM) et les secondes vers les moyens de communication d'adresses de bloc.
- 8Système multiprocesseur selon les revendications 1 et 5 prises ensemble, caractérisé en ce que les moyens de gestion des données partagées comprennent un processeur de gestion mémoire (PGM) associé à la mémoire centrale (RAM) et un processeur de maintien de la cohérence des données partagées (PMC j ) associé à chaque processeur de traitement (CPU j ) et au répertoire de gestion correspondant (RG j ), chaque processeur de maintien de cohérence (PMC j ) étant connecté à un bus de synchronisation (SYNCHRO) piloté par le processeur de gestion mémoire (PGM), de façon à permettre une mise à jour de la mémoire centrale (RAM) et de la mémoire-cache associée (MC j ) en cas de détection d'une adresse de bloc, une mise à jour de la mémoire centrale (RAM) et des mémoires-caches (MC j ) à chaque prélèvement d'adresses dans les registres à décalage complémentaires (RDC j ).
- 9Système multiprocesseur selon les revendications 2 et 5 prises ensemble, caractérisé en ce que les moyens de gestion des données partagées comprennent un processeur de gestion mémoire (PGM) associé à la mémoire centrale (RAM) et un processeur espion de bus (PE j ) associé à chaque processeur de traitement (CPU j ) et au répertoire de gestion correspondant (RG j ), chaque processeur espion de bus (PE j ) et le processeur de gestion mémoire (PGM) étant connectés au bus de communication d'adresses (BUSA) en vue respectivement de surveiller et de traiter les adresses de blocs transmises sur ledit bus de façon à permettre une mise à jour de la mémoire centrale (RAM) et de la mémoire-cache associée (MC j ) en cas de détection d'une adresse de bloc présente dans le répertoire associé (RG j ).
- 10Système multiprocesseur selon les revendications 2 et 5 prises ensemble caractérisé en ce que les moyens de gestion des données partagées comprennent un processeur de gestion mémoire (PGM) associé à la mémoire centrale (RAM) et un processeur de maintien de la cohérence des données partagées (PMC j ) associé à chaque processeur de traitement (CPU j ) et au répertoire de gestion correspondant (RG j ), chaque processeur de maintien de cohérence (PMC j ) étant connecté à un bus de synchronisation (SYNCHRO) piloté par le processeur de gestion mémoire (PGM), de façon à permettre une mise à jour de la mémoire centrale (RAM) et de la mémoire-cache associée (MC j ) en cas de détection d'une adresse de bloc, une mise à jour de la mémoire centrale (RAM) et des mémoires-caches (MC j ) à chaque prélèvement d'adresses sur le bus commun d'adresses BUSA.
- 11Système multiprocesseur selon l'une des revendications 1 à 10, caractérisé en ce que :- plusieurs registres à décalage processeur (RDP k , RDP k+1 ...) correspondant à un ensemble de processeurs déterminé (CPU k , CPU k+1 ...) sont connectés en parallèles à une même liaison série (LS k ), un arbitre local (ABL k ) étant associé à chaque ensemble de processeurs (CPU k , CPU k+1 ...) en vue d'arbitrer les conflits d'accès à la liaison série (LS k ), - un processeur de gestion mémoire (PGM) est relié aux moyens de communication d'adresses de blocs et à la mémoire centrale (RAM) et comprend des moyens de codage adaptés pour associer à chaque bloc d'informations (bi) un entête d'identification du processeur concerné parmi chaque ensemble (CPU k , CPU k+1 ...) partageant une liaison série donnée (LS k ), - les processeurs de gestion (PG k , PG k+1 ...) associés aux mémoires-caches (MC k , MC k+1 ...) des processeurs de l'ensemble précité (CPU k , CPU k+1 ...) comprennent des moyens de décodage de l'en-tête d'identification.
- 12Système multiprocesseur selon l'une des revendications 1 à 11, caractérisé en ce que chaque registre à décalage mémoire (RDM j ) est connecté de façon statique à une liaison série (LS j ) spécifiquement affectée audit registre.
- 13Système multiprocesseur selon l'une des revendications 1 à 11, caractérisé en ce que :- un processeur de gestion mémoire (PGM) est associé à la mémoire centrale (RAM) et comprend une logique (ALLOC) d'affectation des registres à décalage mémoire aux liaisons séries, - les registres à décalage mémoire (RDM₁... RDM j ... RDM n ) sont connectés de façon dynamique aux liaisons séries (LS₁... LS j ...) par l'entremise d'un réseau d'interconnexion (RI) commandé par le processeur de gestion mémoire (PGM).
- 14Système multiprocesseur selon l'une des revendications 1 à 13, dans lequel la mémoire centrale (RAM) est constituée par m bancs mémoires (RAM₁... RAM p ... RAM m ) agencés en parallèle, caractérisé en ce que chaque registre à décalage mémoire (RDM j ) est constitué par m registres à décalage élémentaires (RDM j1 ... RDM jp ... RDM jm ) reliés en parallèles à la liaison série correspondante (LS j ), chaque registre élémentaire (RDM jp ) étant connecté à un banc mémoire (RAM p ) de façon à permettre, en un cycle dudit banc mémoire, un transfert parallèle en lecture ou écriture d'un bloc d'informations (bi) entre ledit registre élémentaire et ledit banc mémoire.
- 15Système multiprocesseur conforme à la revendication 14, dans lequel chaque liaison série (LS j ) est éclatée en m liaisons séries (LS jp ), reliant en point à point chaque processeur (CPU j ) au registre à décalage élémentaire (RDM jp ).
- 16Système multiprocesseur selon la revendication 14, dans lequel chaque banc mémoire (RAM p ) est du type à accès aléatoire doté d'une entrée/sortie de données de largeur correspondant à un bloc d'informations (bi), caractérisé en ce que ladite entrée/sortie de chaque banc mémoire (RAM p ) est reliée par un bus parallèle à l'ensemble des registres élémentaires (RDM 1p ... RDM jp ).
- 17Système multiprocesseur selon l'une des revendications 14, 15 ou 16, synchronisé par une horloge de fréquence F au moins égale à 100 mégahertz, caractérisé en ce que chaque registre à décalage élémentaire mémoire (RDM jp ) et chaque registre à décalage processeur (RDP j ) sont d'un type adapté pour présenter une fréquence de décalage au moins égale à F.
- 18Système multiprocesseur selon l'une des revendications 14, 15 ou 16 synchronisé par une horloge de fréquence F au moins égale à 100 mégahertz, caractérisé en ce que chaque registre à décalage élémentaire mémoire et/ou chaque registre à décalage processeur est constitué d'un ensemble de 2 u sous-registres multiplexés (RDM jp , RDP jp ), chacun apte à présenter une fréquence de décalage au moins égale à F/₂u.
- 19Procédé d'échange d'informations entre une mémoire centrale (RAM) organisée en blocs d'informations (bi) et des processeurs (CPU₁... CPU j ... CPU n ) chacun doté d'une mémoire-cache (MC j ) organisée en blocs de même taille (bi), et d'un répertoire (RG j ) et de son processeur de gestion (PG j ), de façon que l'échange entre mémoire centrale (RAM) et chaque processeur (CPU j ) s'effectue via la mémoire-cache (MC j ) de ce dernier, ledit procédé étant caractérisé en ce que :- chaque transfert de bloc d'informations (bi) depuis la mémoire centrale (RAM) vers la mémoire-cache (MC j ) d'un processeur donné (CPU j ) est effectué à une fréquence de transfert F au moins égale à 100 mégahertz et consiste : . à transférer, en un cycle de mémoire centrale, le bloc (bi) de ladite mémoire centrale (RAM) vers un registre à décalage mémoire (RDM j ) de la taille d'un bloc, faisant partie d'un ensemble de registres à décalage (RDM₁... RDM j ... RDM n ) connectés à la mémoire centrale, . à transférer sur une liaison série (LS j ) le contenu du registre à décalage mémoire (RDM j ) vers un registre à décalage processeur (RDP j ) de même capacité, associé à la mémoire-cache (MC j ) du processeur considéré (CPU j ), . à transférer le contenu dudit registre à décalage processeur (RDP j ) vers ladite mémoire-cache (MC j ), - le transfert de l'adresse d'un bloc d'informations est effectué à la fréquence de transfert F via les liaisons-séries.
- 20Procédé d'échange d'informations entre une mémoire centrale (RAM) organisée en blocs d'informations (bi) et des processeurs (CPU₁... CPU j ... CPU n ) chacun doté d'une mémoire-cache (MC j ) organisée en blocs de même taille (bi), et d'un répertoire (RG j ) et de son processeur de gestion (PG j ), de façon que l'échange entre mémoire centrale (RAM) et chaque processeur (CPU j ) s'effectue via la mémoire-cache (MC j ) de ce dernier, ledit procédé étant caractérisé en ce que :- chaque transfert de bloc d'informations (bi) depuis la mémoire centrale (RAM) vers la mémoire-cache (MC j ) d'un processeur donné (CPU j ) est effectué à une fréquence de transfert F au moins égale à 100 mégahertz et consiste : . à transférer, en un cycle de mémoire centrale, le bloc (bi) de ladite mémoire centrale (RAM) vers un registre à décalage mémoire (RDM j ) de la taille d'un bloc, faisant partie d'un ensemble de registres à décalage (RDM₁... RDM j ... RDM n ) connectés à la mémoire centrale, . à transférer sur une liaison série (LS j ) le contenu du registre à décalage mémoire (RDM j ) vers un registre à décalage processeur (RDP j ) de même capacité, associé à la mémoire-cache (MC j ) du processeur considéré (CPU j ), . à transférer le contenu dudit registre à décalage processeur (RDP j ) vers ladite mémoire-cache (MC j ), - le transfert de l'adresse d'un bloc d'informations est effectué par l'intermédiaire d'un bus commun de communication parallèle d'adresses (BUSA).
- 21Procédé d'échange d'informations entre une mémoire centrale (RAM) organisée en blocs d'informations (bi) et des processeurs (CPU₁... CPU j ... CPU n ) chacun doté d'une mémoire-cache (MC j ) organisée en blocs de mêmes tailles (bi), et d'un répertoire (RG j ) et de son processeur de gestion (PG j ), de façon que l'échange entre mémoire centrale (RAM) et chaque processeur (CPU j ) s'effectue via la mémoire-cache (MC j ) de ce dernier, ledit procédé étant caractérisé en ce que :- chaque transfert de bloc d'informations (bi) depuis la mémoire-cache (MC j ) d'un processeur donné (CPU j ) vers la mémoire centrale (RAM) est effectué à une fréquence de transfert au moins égale à 100 mégahertz et consiste : . à transférer le bloc (bi) de ladite mémoire-cache considérée (MC j ) vers un registre à décalage processeur (RDP j ) de la taille d'un bloc, associé à ladite mémoire-cache (MC j ), . à transférer sur une liaison série (LS j ) le contenu du registre à décalage processeur (RDP j ) vers un registre à décalage mémoire (RDM j ) de même capacité, affecté au processeur considéré dans un ensemble de registres à décalage (RDM₁... RDM j ... RDM n ) connectés à la mémoire centrale (RAM), . à transférer, en un cycle de mémoire centrale, le contenu du registre à décalage mémoire (RDM j ) vers ladite mémoire centrale (RAM), - le transfert de l'adresse d'un bloc d'informations est effectué à la fréquence de transfert F via les liaisons-séries.
- 22Procédé d'échange d'informations entre une mémoire centrale (RAM) organisée en blocs d'informations (bi) et des processeurs (CPU₁... CPU j ... CPU n ) chacun doté d'une mémoire-cache (MC j ) organisée en blocs de mêmes tailles (bi), et d'un répertoire (RG j ) et de son processeur de gestion (PG j ), de façon que l'échange entre mémoire centrale (RAM) et chaque processeur (CPU j ) s'effectue via la mémoire-cache (MC j ) de ce dernier, ledit procédé étant caractérisé en ce que :- chaque transfert de bloc d'informations (bi) depuis la mémoire-cache (MC j ) d'un processeur donné (CPU j ) vers la mémoire centrale (RAM) est effectué à une fréquence de transfert au moins égale à 100 mégahertz et consiste : . à transférer le bloc (bi) de ladite mémoire-cache considérée (MC j ) vers un registre à décalage processeur (RDP j ) de la taille d'un bloc, associé à ladite mémoire-cache (MC j ), . à transférer sur une liaison série (LS j ) le contenu du registre à décalage processeur (RDP j ) vers un registre à décalage mémoire (RDM j ) de même capacité, affecté au processeur considéré dans un ensemble de registres à décalage (RDM₁... RDM j ... RDM n ) connectés à la mémoire centrale (RAM), . à transférer, en un cycle de mémoire centrale, le contenu du registre à décalage mémoire (RDM j ) vers ladite mémoire centrale (RAM), - le transfert de l'adresse d'un bloc d'informations est effectué par l'intermédiaire d'un bus commun de communication parallèle d'adresses (BUSA).
- 23Composant mémoire multiport série, susceptible d'équiper un système multiprocesseur conforme à l'une des revendications 1 à 18, caractérisé en ce qu'il est constitué par un circuit intégré comprenant une mémoire à accès aléatoire (RAM) de largeur prédéterminée correspondant à un bloc d'informations (bi), un ensemble de registres à décalage (RDM₁... RDM j ... RDM n ), chacun de capacité correspondant à la largeur de la mémoire, un bus parallèle interne (BUSI) reliant l'accès de la mémoire et les registres à décalage, une logique de sélection d'un registre à décalage (LSR) adaptée pour valider la liaison sur le bus interne entre la mémoire et un registre à décalage prédéterminé, et un ensemble de broches externes d'entrée/sortie (adbloc, admot, numreg, cs, wr, rd, bitbloc, normal/config, hi, di) pour l'entrée d'adresses vers la mémoire (RAM), pour l'entrée d'adresses vers la logique de sélection (LSR), pour l'entrée et la validation de commandes de transfert en lecture ou écriture d'un bloc d'informations (bi) entre la mémoire (RAM) et les registres à décalage (RDM j ), pour l'entrée d'un signal d'horloge vers chaque registre à décalage (RDM j ), pour l'entrée bit à bit d'un bloc d'informations (bi) vers chaque registre à décalage (RDM j ) et pour la sortie bit à bit d'un bloc d'informations de chaque registre à décalage (RDM j ).
- 24Composant selon la revendication 23, caractérisé en ce qu'il comprend au moins un registre de configuration (RC₁, RC₂...) possédant des entrées de programmation, chaque registre de configuration étant relié à une logique de forçage (LF) connectée à la mémoire (RAM) et aux registres à décalage (RDM j ) en vue d'assurer le forçage d'états de ladite mémoire et desdits registres à décalage.
- 25Composant selon la revendication 24, permettant le choix de la taille des blocs d'informations (bi) traités, caractérisé en ce que :- la mémoire (RAM) est découpée en zones combinables pour permettre la mémorisation des diverses tailles possibles de blocs d'informations, - chaque registre à décalage (RDM j ) est découpé en tronçons combinables pour permettre de charger les diverses tailles possibles de blocs d'informations, avec des dérivations aptes à assurer le décalage correspondant à chaque taille, - le bus interne (BUSI) est doté d'une logique de multiplexage (MT) pour permettre les transferts de blocs d'informations (bi) des diverses tailles entre les combinaisons de zones de la mémoire (RAM) et les combinaisons correspondantes de tronçons des registres à décalage (RDM j ), - un registre de configuration (RC₁) est prévu de capacité correspondant au nombre de tailles de blocs possibles, - la logique de forçage (LF) reliée au registre (RC₁ ) comprend une unité logique adaptée pour commander la logique de multiplexage (MT) en vue de valider les transferts de blocs d'informations (bi) dans une taille donnée correspondant au paramètre contenu dans le registre de configuration (RC₁).
- 26Composant selon l'une des revendications 24 ou 25, caractérisé en ce que :- l'entrée et la sortie de chaque registre à décalage (RDM j ) sont reliées à une même broche externe par l'intermédiaire d'une porte logique (PL j ), - un registre de configuration (RC₂) est prévu de capacité correspondant au nombre de registres à décalage (RDM j ), - la logique de forçage (LF) reliée au registre de configuration (RC₂) comprend une unité logique adaptée pour commander les portes logiques (PL j ) en vue de forcer le fonctionnement de chaque registre à décalage (RDM j ) en mode entrée ou en mode sortie en fonction d'un bit contenu dans le registre de configuration (RC₂) affecté audit registre à décalage (RDM j ).
- 27Composant selon l'une des revendications 24, 25 ou 26, caractérisé en ce que :- l'entrée et la sortie de chaque registre à décalage (RDM j ) sont reliées à une même broche externe par l'intermédiaire d'une porte logique (PL j ), - un registre de configuration (RC₃) est prévu de capacité correspondant au nombre de registres à décalage (RDM j ), - la logique de forçage (LF) reliée au registre de configuration (RC₃) comprend une unité logique reliée à la commande de lecture de la mémoire (RAM) et adaptée pour commander chaque porte logique (PL j ) soit en mode sortie au moment de la lecture de la mémoire (transfert de la mémoire RAM vers le registre correspondant RDM j ) pendant toute la durée du vidage dudit registre à décalage (RDM j ), soit en mode entrée, le reste du temps.
- 28Composant selon l'une des revendications 24, 25, 26 ou 27, caractérisé en ce qu'il comprend une broche externe d'entrée (bit/bloc), une ou des broches externes d'entrée/sortie d'informations (data), une logique de commande (COM) reliée à la broche d'entrée (bit/bloc), aux broches d'entrée/sortie (data), à la mémoire (RAM) et à la logique de sélection (LSR) et adaptée selon l'état de l'entrée (bit/bloc) pour engendrer soit les transferts de blocs d'informations (bi) entre mémoire (RAM) et registres à décalage (RDM j ), soit des transferts de bits directement entre la mémoire (RAM) et les broches (data).
- 29Composant selon les revendications 24 et 28 prises ensemble, caractérisé en ce que les registres de configuration (RC₁, RC₂...) sont reliés :- d'une part à la logique de sélection (LSR) laquelle est adaptée pour sélectionner lesdits registres de configuration pour des adresses prédéterminées affectées auxdits registres, - d'autre part, à la logique de commande (COM) laquelle est adaptée pour transmettre les données en provenance des broches d'entrée/sortie (data) vers lesdits registres de configuration en vue de leur programmation.
- 30Composant selon l'une des revendications 23 à 29, dans lequel sur le bus interne (BUSI) reliant l'accès de la mémoire (RAM) et les registres à décalage (RDM j ) est interposée une logique de type "barillet" (BS) ("barrel shifter") apte à assurer une permutation circulaire sur les bits de chaque bloc d'informations, ladite logique (BS) possédant une entrée de commande du pas de glissement en unité mot connectée à des broches d'entrée (admot).
Independent claims30
284 paragraphs, as filed
The invention relates to a multiprocessor type system comprising a central memory, processing processors and cache memories associated with the processing processors. It extends to a process for exchanging information between central memory and processing processors via the cache memory associated with each of these processors. It also targets a new integrated circuit component, capable of equipping the multiprocessor system.
We know that, in the most common known multiprocessor systems, all the information (data, addresses) pass through a common parallel communication bus between the central memory and the various processing processors, which constitutes a bottleneck: its throughput is indeed insufficient to supply all the processors at full output, from a common central memory.
To increase the speed of information transfer, a first solution consists in associating with each processing processor a cache memory which, by the locality of the information, makes it possible to reduce the requests towards the central memory. However, in the case where the volume of data shared between processors is substantial, maintaining the consistency of the data between memories generates additional information traffic on the communication bus which is opposed to a significant reduction in the overall speed on this bus and, therefore, removes much of its interest in this solution.
Another solution consists in making the communication bus in the form of a mesh network designated by "crossbar", which allows direct communication between each processing processor and each subset of the central memory (memory bank). However, this solution is very cumbersome and very expensive to produce due to the very high number of interconnections, and it becomes completely unrealistic beyond a dozen processing processors. In addition, in the event of multiple requests from several processors on the same memory bank, such a solution involves access conflicts, a source of slowing down of exchanges.
Another more common solution due to its architectural simplicity consists in associating a local memory with each processing processor to store data specific to it, and in memorizing the shared data in the common central memory. However, the big defect of this architecture is its non-transparency, that is to say the need for the programmer to organize the detail of the data assignments in the various memories, so that this solution is of very use binding. In addition, in the event of a high volume of shared data, it can lead, as before, to saturation of the access bus to the central memory.
In addition, a solution called "aquarius architecture" has been proposed by the University of Berkeley and consists in improving the aforementioned crossbar solution by combining with the crossbar network, on the one hand, for the non-shared data, cache memories which are connected to the crossbar network, on the other hand, for shared data, separate cache memories which are connected to a common synchronization bus. This solution brings a gain in speed of exchanges but remains very cumbersome and very expensive to carry out.
The present invention proposes to provide a new solution, making it possible to considerably increase the information exchange rates, while keeping a transparent architecture for the user, much simpler than the crossbar architecture.
An objective of the invention is thus to make it possible to significantly increase the number of processors in the system, while benefiting from a high efficiency for each processor.
Another objective is to provide an integrated circuit component structure, allowing a very simple implementation of the architecture of this new multiprocessor system.
To this end, the multiprocessor system targeted by the invention is of the type comprising a central memory (RAM) organized in information blocks (bi), processing processors (CPU₁ ... CPU<sub>j</sub>... CPU<sub>not</sub>), a cache memory (MC<sub>j</sub>) linked to each processing processor (CPU<sub>j</sub>) and organized in information blocks (bi) of the same size as those in the main memory, a directory (RG<sub>j</sub>) and its management processor (PG<sub>j</sub>) associated with each cache memory (MC<sub>j</sub>), means of communication of block addresses between processors (CPU<sub>j</sub>) and main memory (RAM); according to the present invention, said multiprocessor system is provided with:<ul id="ul0001" list-style="none"><li>. a set of shift registers, called memory shift registers (RDM₁ ... RDM<sub>j</sub>... RDM<sub>not</sub>), each register (RDM<sub>j</sub>) of this assembly being connected to the central memory (RAM) so as to allow, in one cycle of this memory, a parallel transfer in reading or writing of a block of information (bi) between said register and said central memory,</li><li>. shift registers, called processor shift registers (RDP₁ ... RDP<sub>j</sub>... RDP<sub>not</sub>), each processor shift register (RDP<sub>j</sub>) being connected to the cache memory (MC<sub>j</sub>) a processor (CPU<sub>j</sub>) so as to allow a parallel transfer in read or write of an information block (bi) between said shift register (RDP)<sub>j</sub>) and said cache memory (MC<sub>j</sub>),</li><li>. a set of serial links (LS₁ ... LS<sub>j</sub>... LS<sub>not</sub>), each connecting a memory shift register (RDM<sub>j</sub>) and a processor shift register (RDP<sub>j</sub>) and adapted to allow the transfer of information blocks (bi) between the two registers considered (RDM<sub>j</sub>, RDP<sub>j</sub>).</li></ul>
Thus, in the multiprocessor system according to the invention, the exchanges between cache memories and associated processors are carried out as in conventional systems equipped with cache memories. On the other hand, exchanges between central memory and cache memories take place in an entirely original way.
Each transfer of information block (bi) from the central memory (RAM) to the cache memory (MC<sub>j</sub>) of a given processor (CPU<sub>j</sub>) consists of:<ul id="ul0002" list-style="none"><li>. to transfer, in a central memory cycle, the block (bi) from said central memory (RAM) to the memory shift register (RDM<sub>j</sub>) (the size of a block) which is directly connected to the main memory and which corresponds to the processor (CPU<sub>j</sub>) considered,</li><li>. to be transferred to the corresponding serial link (LS<sub>j</sub>) the contents of this memory shift register (RDM<sub>j</sub>) to the processor shift register (RDP<sub>j</sub>) (of the same capacity) which is associated with the cache memory (MC<sub>j</sub>) of the processor considered (CPU<sub>j</sub>),</li><li>. to transfer the content of said processor shift register (RDP<sub>j</sub>) to said cache memory (MC<sub>j</sub>).</li></ul>
In the opposite direction, each transfer of information block (bi) from the cache memory (MC<sub>j</sub>) of a given processor (CPU<sub>j</sub>) to the main memory (RAM) consists of:<ul id="ul0003" list-style="none"><li>. to transfer the block (bi) from said cache memory considered (MC<sub>j</sub>) to the processor shift register (RDP<sub>j</sub>) which is associated with said cache memory (MC<sub>j</sub>),</li><li>. to be transferred to the corresponding serial link (LS<sub>j</sub>) the content of the processor shift register (RDP<sub>j</sub>) to the memory shift register (RDM<sub>j</sub>), assigned to the processor considered (among the set of shift registers (RDM₁ ... RDM<sub>j</sub>... RDM<sub>not</sub>) connected to the central memory (RAM)),</li><li>. to transfer, in a central memory cycle, the content of the memory shift register (RDM<sub>j</sub>) to said central memory (RAM).</li></ul>
Under these conditions, the transfer of each block of information (bi) is carried out, no longer through a parallel bus as is the case in known systems, but by high speed serial links. These serial links make it possible to obtain transfer times of each block (bi) comparable and even lower than the transfer times in known systems with parallel bus. The comparative example given below with common parameter values for current technology clearly illustrates this seemingly paradoxical fact.
It is assumed that each block of information (bi) is of a size equal to 64 bytes.
In the system of the invention, the transfer time between the central memory and a cache memory is broken down into:<ul id="ul0004" list-style="dash"><li>a central memory (RAM) / memory shift register (RDM) transfer time<sub>j</sub>): 100 nanoseconds (performance of a current type random access central memory),</li><li>a serial transfer time on the corresponding serial link: 64 x 8 x 1 / 500.10⁶, or 1,024 nanoseconds, assuming a transfer frequency of 500 Megahertz (not exceptional with current technologies which allow frequencies reaching 3,000 Megahertz ),</li><li>a processor shift register transfer time (RDP)<sub>j</sub>) / cache (MC<sub>j</sub>): 50 nanoseconds (very common type cache memory).</li></ul>
The total transfer time of a block is therefore of the order of 1,200 nanoseconds (by integrating second-order sequence delays).
In known systems with cache memories in which the exchange of information takes place directly in parallel by 4-byte words (the most common systems lead to buses of the usual type with 32 data wires), the transfer time d 'a block is equal to the transfer time of the 16 words of 4 bytes which constitute this block, that is to say: 16 x 100 = 1600 nanoseconds.
We therefore see that, with average hypotheses in the two solutions, these times are comparable. However, if we compare the architecture of the system in accordance with the invention with that of a common parallel bus with cache memories (first solution mentioned above), we see that:<ul id="ul0005" list-style="none"><li>. in the conventional solution (parallel common bus), the central memory and the common bus are 100% occupied during the transfer, since information circulates between the two throughout the duration of the transfer,</li><li>. in the system according to the invention, the serial link is occupied at 100% during the transfer, but the main memory is occupied less than 10% of the transfer time (memory read time and loading of the memory shift register (RDM)<sub>j</sub>)), so that the central memory can serve 10 times more processors than in the previous case (the occupation of the serial link being of no importance since it is private and assigned to the processor).</li></ul>
It should be emphasized, moreover, that in the system of the invention, each serial link which links each processor individually to the central memory is a single link (with one or two data wires), so that the serial network thus constituted is not comparable in terms of complexity with, for example, a crossbar network, each link of which is a parallel link with a multiplicity of wires (32 data wires in the aforementioned comparative example), with all the necessary switches.
In addition, as will be seen below on the comparative curves, the system according to the invention has very improved performance compared to traditional common bus systems and in practice allows the implementation of a number of processing processors. much higher (from several tens to a hundred processors); these performances are compatible with those of a crossbar system, but the system according to the invention is of much greater architectural simplicity.
In the system of the invention, each serial link can in practice be produced either by means of two unidirectional serial links for bit-to-bit transfer, or by means of a single bidirectional link.
In the first case, each memory shift register (RDM<sub>j</sub>) and each processor shift register (RDP<sub>j</sub>) are split into two registers, one specialized for transfer in one direction, the other for transfer in the other direction. The two unidirectional serial links are then connected to the memory shift register (RDM<sub>j</sub>) split and to the corresponding processor shift register (RDP<sub>j</sub>) split, so as to allow, one a transfer in one direction, the other a transfer in the other direction.
This embodiment with two unidirectional links has the advantage of not requiring any transfer management on a link, but the disadvantage of doubling the necessary resources (link, registers).
In the second case, a logic of validation of the direction of transfer is associated with the bidirectional link so as to allow an alternate transfer in both directions on said link. This logic can be integrated into the management processor (PG<sub>j</sub>) associated with the cache memory (MC<sub>j</sub>) to which is linked said bidirectional link.
Of course, each serial link can possibly be produced with a higher number of serial links.
In the multiprocessor system according to the invention, the address communication means can take essentially two embodiments: first, they can consist of a block address parallel communication bus (BUSA), common to all processors (CPU<sub>j</sub>) and connecting these and the central memory (RAM) with a conventional bus arbiter (AB) adapted to manage access conflicts to said bus. Note that this address bus is only used for the communication of block addresses: in terms of structure, this bus is identical to the parallel address communication bus of known systems, for which no saturation problem, since it can be released immediately after transfer of the block address.
However, another embodiment of these address communication means can be envisaged in the multiprocessor system of the invention, consisting in using the serial transfer links of the information blocks (bi) to transfer the addresses of these blocks .
In this case, a complementary shift register (RDC<sub>j</sub>) is connected on each serial link (LS<sub>j</sub>) in parallel with the corresponding memory shift register (RDM<sub>j</sub>): the addresses transmitted by said serial link are thus loaded into each of these additional registers (RDC<sub>j</sub>); an access management arbitrator (ABM) linked to said registers (RDC)<sub>j</sub>) and to the central memory (RAM) is then provided for taking the addresses contained in said registers and for managing conflicts of access to the central memory (RAM). Such an arbitrator is of known design in itself, this type of access conflict has now been resolved for many years. In this embodiment, the presence of a parallel address communication bus is avoided, but the management resources are increased.
Furthermore, the multiprocessor system according to the invention is particularly well suited to efficiently manage the problems of consistency of the data shared between processing processors. Indeed, the conventional solutions for managing this shared data find their limit in known systems due to the bottleneck in the communication of information, but on the contrary become perfectly satisfactory and efficient in the system of the invention where a such a bottleneck no longer exists, so that this system can be equipped with shared data management means of design similar to those of known systems.
For example, a traditional solution for managing shared data consists in preventing it from passing through the cache memories: conventionally, partition logic (LP<sub>j</sub>) is associated with each processing processor (CPU<sub>j</sub>) in order to differentiate the addresses of the shared data and those of the non-shared data so as to direct the former directly to the central memory (RAM) and the latter to the corresponding cache memory (MC<sub>j</sub>).
In a first version of the architecture according to the invention, the system comprises:<ul id="ul0006" list-style="none"><li>. a special parallel word communication bus (BUSD) connecting the processors (CPU<sub>j</sub>) and the main memory (RAM),</li><li>. partition logic (LP<sub>j</sub>), associated with each processor (CPU<sub>j</sub>) and adapted to differentiate the addresses of the shared data and those of the non-shared data so as to transmit these on the means of address communication with their identification,</li><li>. decoding logic (DEC) associated with the central memory (RAM) and suitable for receiving addresses with their identification and routing the data at memory output either to the corresponding memory shift register (RDM<sub>j</sub>) for non-shared data, or to the special word communication bus (BUSD) for shared data.</li></ul>
This solution has the advantage of being very simple architecturally; the presence of the special parallel communication bus (BUSD) leads to better performance compared to a solution which would consist in using the serial links to transfer not only the blocks of unshared data but also the words of shared data. It should be noted that this last solution can, if necessary, be envisaged in the event of a low rate of shared data.
In another version, the system is provided with a special bus for parallel word communication (BUSD) and a special common bus for word address communication (BUSAM), connecting the processors (CPU<sub>j</sub>) and main memory (RAM). The partition logic (LP<sub>j</sub>) switches the addresses of the shared data to the common special bus (BUSAM) for the transfer of data by the special word bus (BUSD), and switches the unshared data to the address communication means (such as consist of a parallel communication bus or that communication takes place via serial links).
The presence of a special word address communication bus makes it possible, in this version, to reduce the saturation limit of the address communication means, in the event of high requests for shared data.
Another version which will be preferred in practice in the case where the address communication means are constituted by a parallel address communication bus (BUSA), consists in equipping the system with an associated memory management processor (PGM). memory (RAM) and a spy bus processor (PE<sub>j</sub>) associated with each processing processor (CPU<sub>j</sub>) and to the corresponding management directory (RG<sub>j</sub>); the memory management processor (PGM) and each spy bus processor (PE<sub>j</sub>), of structures known per se, are connected to the address communication bus (BUSA) respectively for the purpose of monitoring and processing the block addresses transmitted on said bus so as to allow updating of the central memory (RAM ) and the associated cache memory (MC<sub>j</sub>) if a block address found in the associated directory is detected (RG<sub>j</sub>).
The memory management processor (PGM) and each spy processor (PE<sub>j</sub>) associate status bits with each block of information, update them according to the nature (read or write) of block requests which pass on the bus (BUSA) and ensure the consistency of the shared data using these status bits which allow them to force or not the writing of a block in central memory at the time of the requests on the bus (BUSA).
In the case mentioned above where the address communications are made by the serial links, the management of the shared data can also be ensured in a centralized manner, by a memory management processor (PGM) associated with the central memory (RAM) and a shared data consistency processor (PMC)<sub>j</sub>) associated with each processing processor (CPU<sub>j</sub>) and to the corresponding management directory (RG<sub>j</sub>), each consistency maintaining processor (PMC<sub>j</sub>) being connected to a synchronization bus (SYNCHRO) controlled by the memory management processor (PGM), so as to allow updating of the central memory (RAM) and the associated cache memory (MC<sub>j</sub>) if a block address is detected, an update of the central memory (RAM) and the cache memories (MC<sub>j</sub>) each time addresses are removed from the additional shift registers (RDC)<sub>j</sub>).
As before, this update is ensured by means of status bits associated with each information block by the processor (PGM).
It should be noted that a synchronization bus of the type defined above can if necessary be provided on the previous architecture where the block addresses transit on a common bus of BUSA addresses. In this case, the spy processors (PE<sub>j</sub>) are requested by the memory management processor (PGM) via the synchronization bus, and this only when they are concerned by the transfer. This avoids unnecessary access to the cache memories; the spy processors then become passive (since they are solicited by the PGM processor) and we will rather designate them by the more appropriate expression of "coherence maintaining processor" according to the terminology used above.
Another solution is to reserve the parallel address communication bus (BUSA) for the transfer of addresses of shared data blocks and to use the serial links for the transfer of unshared data blocks.
Furthermore, the multiprocessor system according to the invention lends itself to grouping of processors on the same serial link, so as to limit the serial links and the corresponding memory shift registers (RDM<sub>j</sub>) required.
The number of memory shift registers (RDM<sub>j</sub>) can correspond to the number of serial links (LS<sub>j</sub>), in which case each memory shift register (RDM<sub>j</sub>) is statically connected to a serial link (LS<sub>j</sub>) specifically assigned to said register.
The number of memory shift registers (RDM<sub>j</sub>) may also be different from that of serial links (LS<sub>j</sub>) and in particular lower, in which case these registers are dynamically connected to the serial links (LS<sub>j</sub>) through an interconnection network.
As in conventional systems, the central memory (RAM) can be divided into m memory banks (RAM₁ ... RAM<sub>p</sub>... RAM<sub>m</sub>) arranged in parallel. Each memory shift register (RDM<sub>j</sub>) is then made up of m elementary registers (RDM<sub>d1</sub>... RDM<sub>jp</sub>... RDM<sub>jm</sub>) connected in parallel to the corresponding serial link (LS<sub>j</sub>). However, an additional level of parallelism and better electrical or optical adaptation of the link are obtained in a variant where each RAM memory bank<sub>p</sub> is connected to each CPU processor<sub>j</sub> via a point-to-point serial link LS<sub>jp</sub>.
Furthermore, to present transfer performance at least equal to that of conventional parallel bus systems, the system according to the invention is preferably synchronized by a clock of frequency F at least equal to 100 megahertz. Memory shift registers (RDM<sub>j</sub>) and processor shift registers (RDP<sub>j</sub>) can very simply be of a type suitable for exhibiting an offset frequency at least equal to F.
In the case of very high frequencies (notably greater than 500 megahertz with current technology), these registers can be divided into sub-registers of lower offset frequency, and multiplexed.
The invention extends to a serial multiport memory component, capable of equipping the previously defined multiprocessor system, with a view to simplifying its manufacture. This component, which may also have different applications, consists of an integrated circuit comprising a random access memory (RAM) of predetermined width corresponding to an information block (bi), a set of shift registers (RDM₁. .. RDM<sub>j</sub>... RDM<sub>not</sub>), each of capacity corresponding to the width of the memory, an internal parallel bus (BUSI) connects the memory access and the shift registers, a logic for selecting a shift register (LSR) adapted to validate the link on the internal bus between the memory and a predetermined shift register, and a set of external input / output pins for the address input to the memory (RAM), for the address input to the logic selection (LSR), for entering and validating transfer commands for reading or writing a block of information (bi) between memory (RAM) and shift registers (RDM)<sub>j</sub>), for inputting a clock signal to each shift register (RDM<sub>j</sub>), for the bit by bit input of an information block (bi) to each shift register (RDM<sub>j</sub>) and for the bit-by-bit output of an information block from each shift register (RDM<sub>j</sub>).
This component can be made configurable by adding configuration registers (RC₁, RC₂, ...) allowing in particular a choice of size of the information blocks (bi) and of various operating modes of the shift registers.
The invention having been explained in its general form is illustrated by the following description with reference to the accompanying drawings which present, without limitation, several embodiments; on these drawings which form an integral part of this description:<ul id="ul0007" list-style="dash"><li>FIG. 1 is a block diagram of a first embodiment of the multiprocessor system according to the invention,</li><li>FIG. 2 is a diagram giving the calculated performance curve of this system (A) and, for comparison, the corresponding curve (B) for a conventional multiprocessor architecture with common bus,</li><li>FIGS. 3, 4 and 5 are detailed logic diagrams of functional units of the system of FIG. 1,</li><li>FIG. 6 is a block diagram of another embodiment of the system,</li><li>FIG. 7 is a block diagram of a system of the type of that of FIG. 1, provided with means for managing shared data,</li><li>FIG. 8 is a detailed logic diagram of a subset of the system of FIG. 7,</li><li>FIG. 9 is a block diagram of a system similar to that of FIG. 7 with a variant of shared data management means,</li><li>FIG. 10 is a block diagram of a similar system, provided with different shared data management means,</li><li>FIGS. 11, 12a, 12b, 12c, 12d, 13, 14, 15, 16, 17 are detailed logic diagrams of functional units of the processor system of FIG. 10,</li><li>FIG. 18 is a block diagram of a system of the type of that of FIG. 6, provided with means for managing shared data,</li><li>FIG. 19 is a simplified block diagram of a variant of the system, in which several central units share the same serial link,</li><li>FIG. 20a is a block diagram of a preferred embodiment, in which the central memory is organized into several memory banks,</li><li>FIG. 20b is a variant of the architecture shown in FIG. 20a,</li><li>FIGS. 21a and 21b show schematically another structure of RAM memory capable of equipping said system,</li><li>FIG. 22 is a block diagram showing the structure of a serial multiport memory component, capable of equipping the system.</li></ul>
The device presented in the form of a block diagram in FIG. 1 is a multiprocessor system having n processing processors CPU₁ ... CPU<sub>j</sub>... CPU<sub>not</sub>. This figure shows two processing processors CPU₁ and CPU<sub>j</sub> with their associated logic. Each of these processing processors is of a traditional type, for example "MOTOROLA 68020" or "INTEL 80386" ... and can include local resources in memories and peripheral interfaces and be equipped with a virtual memory device.
The device comprises a central memory with random access RAM produced in a conventional manner from integrated memory circuits: in particular dynamic RAM "INTEL" "NEC" "TOSHIBA" ... of 256 Kbits, 1 Mbits, 4 Mbits ... according to the application. This memory is organized into information blocks bo ... bi ... of determined size t (usually 256 bits at 2 Kbits) and the access edge of said memory corresponds to the size of a block.
The main memory is connected in parallel to n shift registers RDM₁ ... RDM<sub>j</sub>... RDM<sub>not</sub> said memory registers, each memory register having the size t of an information block; each of these registers is produced using very fast technology ("ASGA"), a block that can be loaded or unloaded in one cycle from the central memory RAM. The number n of registers is equal to the number of CPU processors<sub>j</sub>.
In addition, an MC cache memory<sub>j</sub> is associated in a manner known per se with each CPU processor<sub>j</sub> ; each cache memory is conventionally constituted by a fast memory, with random access, of small capacity compared to the central memory RAM. An RG directory<sub>j</sub> and a PG management processor<sub>j</sub> are traditionally connected to the cache memory and to the processing processor to manage the information passing through the cache memory.
Furthermore, in the system of the invention, an RDP shift register<sub>j</sub> said register-processor, is connected by its parallel port, to each MC cache memory<sub>j</sub> ; each RDP processor register<sub>j</sub> is of size corresponding to that of a bi block and of structure similar to that of the RDM memory registers<sub>j</sub>.
Each RDM memory register<sub>j</sub> is connected by its serial port, to the serial port of an RDP register-processor<sub>j</sub> by a serial link LS<sub>j</sub>. Examples of embodiment of this serial link which can include a bidirectional link or two unidirectional links are illustrated in FIGS. 4 and 5. The control of the transfer of the blocks bi between corresponding registers RDM<sub>j</sub> and RDP<sub>j</sub> is provided by TFR transfer logic<sub>j</sub> and TFR '<sub>j</sub> which are symmetrically associated with the RDM memory register<sub>j</sub> and to the RDP processor register<sub>j</sub> ; an exemplary embodiment of these transfer logics (in themselves conventional) is detailed in FIG. 3.
The central memory unit RAM, memory shift registers RDM₁ ... RDM<sub>not</sub> and associated transfer logic TFR₁ ... TFR<sub>not</sub> constitute a functional assembly called "serial multiport memory" MMS. The CPU processing processor assembly<sub>j</sub>, MC cache<sub>j</sub>, RG cache management directory<sub>j</sub>, PG cache management processor<sub>j</sub>, RDP processor shift register<sub>j</sub> and associated transfer logic TFR '<sub>j</sub> constitutes a "functional" unit called "central unit" CPU<sub>j</sub>.
Furthermore, the system comprises means for communicating addresses of blocks of CPU processors.<sub>j</sub> to the central memory RAM, constituted in the example by a common parallel communication bus BUSA on which the CPUs are connected<sub>j</sub> (through their PG management processor<sub>j</sub>) and the central memory RAM.
Access to the BUSA bus is traditionally regulated by an AB bus arbitrator.
The general operation of the architecture defined above is as follows: CPU processor<sub>j</sub> executes its own program consisting of instructions, which are in the form of words in central memory RAM with extracts in the associated cache memory MC<sub>j</sub>. On program instructions, the CPU processor<sub>j</sub> is led either to read data words which themselves are in the central memory RAM or in the cache memory MC<sub>j</sub> in the form of extracts, either to write data words in the central memory RAM and in the cache memory MC<sub>j</sub>.
Each CPU operation<sub>j</sub> (designated "request") requires the supply of the address adr of the word concerned, the nature r, w of the operation (read, write) and the exchange (data) of the word concerned.
Each word request activates the PG processor<sub>j</sub> which then consults in a classic way the directory of the RG cache<sub>j</sub> which indicates whether the block bi containing the word concerned is present in the cache memory MC<sub>j</sub> and, if necessary, the block frame in the cache memory where the sought block is located.
If the block bi containing the word concerned is in the cache memory MC<sub>j</sub>, then when read this word is read in said cache memory and sent to the processor CPU<sub>j</sub> ; when writing the word provided by the CPU<sub>j</sub> is written in the cache memory: the memory transaction is complete.
If the block containing the word concerned is not in the MC cache memory<sub>j</sub>, then a reading of block bi in central memory RAM is necessary. Two cases can occur.
First case
MC cache memory<sub>j</sub> has at least one free block location, determined by the PG processor<sub>j</sub> using status bits associated with each entry in the RG directory<sub>j</sub>. In this case, the PG processor<sub>j</sub> typically requires the BUSA bus by submitting its request to the bus arbiter AB. The latter in turn grants the BUSA bus to the PG processor<sub>j</sub> which has read access to the central memory RAM, the block read from memory being loaded into the register RDM<sub>j</sub>, identified by the origin j of the call. The end of the reading cycle results in the release of the BUSA bus and the activation of the transfer with the serial link LS<sub>j</sub> allow to transfer the contents of the RDM memory register<sub>j</sub> in the RDP processor register<sub>j</sub>. End of transfer activates writing to MC cache memory<sub>j</sub> the contents of the processor register in the block space reserved for this purpose and the transaction can end as before.
Second case
MC cache memory<sub>j</sub> does not have a free space, then, by a conventional algorithm, a cache space is made a candidate to receive the requested block. Two situations can be encountered:<ul id="ul0008" list-style="dash"><li>the block contained in the candidate location has not been modified since its installation: it is simply eliminated by freeing the block frame by a simple writing of a status bit in the directory (RG<sub>j</sub>) and the transaction can continue as before,</li><li>the block contained in the candidate location has been modified and an update of the central memory RAM is necessary. To do this, the management processor PG<sub>j</sub> transfers the candidate block to the RDP processor register<sub>j</sub>, activates the transfer of the RDP processor register<sub>j</sub> to the RDM memory register<sub>j</sub>, then requests the common bus BUSA by submitting its request to the referee AB. When the arbitrator grants the bus to the management processor PG<sub>j</sub>, the latter activates a write command which has the effect of transferring the content of the memory register RDM<sub>j</sub> to its location in central RAM memory. The RAM memory update is complete and the transaction can continue as before.</li></ul>
Thus, in the device of the invention, the exchanges between the processing processors CPU<sub>j</sub> and their MC cache<sub>j</sub> and associated logic RG<sub>j</sub> , PG<sub>j</sub> , are carried out in a conventional manner; on the other hand, block transfers between central RAM memory and MC cache memory<sub>j</sub> no longer pass by a common parallel bus but by LS serial links<sub>j</sub> dedicated to each CPU processing processor<sub>j</sub>, the common bus BUSA only serving for the transfer of addresses and thus having considerably reduced traffic.
We know that, for classic common bus architectures, a modeling studied by "PATEL" (analysis of multiprocessors with private cache, JANAK H. PATEL - IEEE Transactions on computers vol. C.31, N ° 4 APRIL 1982) to the following approximate formula giving the yield U as a function of the number of processors present:<maths id="math0001" num=""><img file="EP0346420B1_D0001.tif" /></maths> or<dl id="dl0001"><dt>the yield U</dt><dd>is the average utilization rate of each processing processor,</dd><dt>m</dt><dd>is the probability for a processing processor to make a memory request, not present in its cache memory (this probability m = α .Pa is proportional to the probability of absence Pa of the information in the cache memory and to a factor α depending on the power of the processing processor reduced to a percentage of memory requests),</dd><dt>W</dt><dd>is the average waiting time of the common bus, which is a function of the number n of processors,</dd><dt>tf</dt><dd>is the transfer time of a block from central memory to a cache memory.</dd></dl>
The hypotheses on the basis of which this formula was established show that it is applicable to the architecture in accordance with the invention, with a level of approximation comparable to the level of approximation of the formula for conventional common bus architectures.
It is thus possible to compare the performance of the two types of architecture by assuming that the components common to the two architectures have identical characteristics.
FIG. 2 gives the curves obtained from the yield U as a function of the number n of processors for the following parameters, the parameters common to the two devices being identical, and all of usual value:<ul id="ul0009" list-style="dash"><li>block size bi = 64 bytes,</li><li>word size for parallel transfer on common bus = 4 bytes,</li><li>RAM central memory access time = 100 nanoseconds,</li><li>BUSA bus cycle time = 50 nanoseconds,</li><li>serial transfer frequency = 500 MHz,</li><li>probability of absence Pa = 0.005 (cache memory of 16 Kbytes),</li><li>processor power factor: α = 0.5.</li></ul>
It can be seen by comparing the curves A (architecture of the invention) and B (classical architecture) that the architecture according to the invention has a clearly higher efficiency than the conventional architecture; the architecture of the invention makes it possible to set up a number of processors much higher than conventional common bus architectures which in practice cannot exceed ten processors. For example, in the classic case, a yield of 0.75 is obtained from the tenth processor, while it is obtained for more than 80 processors in the case of the invention.
Figure 3 shows an embodiment of a TFR transfer logic<sub>j</sub> or TFR '<sub>j</sub> allowing to transfer a bi block of information from an RDM memory register<sub>j</sub> to an RDP processor register<sub>j</sub> (the reverse transfer is ensured by symmetrical means not shown in this figure). Each TFR logic<sub>j</sub> or TFR '<sub>j</sub> includes a TFRE broadcast control part<sub>j</sub> and TFRE '<sub>j</sub> and a part of TFRR reception control<sub>j</sub> and TFRR '<sub>j</sub> which are cross-activated (TFRE broadcast<sub>j</sub> activated in synchronism with TFRR reception '<sub>j</sub>). The system includes a clock generator H whose frequency fixes the transmission speed and supplies the clock signal h to the transmission part TFRE<sub>j</sub> and at the reception part TFRR '<sub>j</sub>.
In the TFRE broadcast section<sub>j</sub> a DC downcounter register receiving by its loading input <o ostyle="single">load₂</o> the read signal <o ostyle="single">r</o> of the PG management processor<sub>j</sub> allows t + 1 clock pulses h to pass through an ET1 logic gate controlled by a zero crossing signal <o ostyle="single">"borrow"</o>, the output of this gate ET1 being connected to the down counting down input of the DC down counter and to the shift1 shift input of the RDM memory register<sub>j</sub>.
In the TFRR reception section '<sub>j</sub>, a flip-flop B is connected by its data input D to the serial output of the register-processor RDP<sub>j</sub>, the clock input clk of this flip-flop being connected to the clock H to receive the signal h. An initialization signal "<o ostyle="single">init</o>"provided by the PG management processor<sub>j</sub> is connected to the entrance <o ostyle="single">S</o> of scale B and at the loading entrance <o ostyle="single">load3</o> the RDP processor register<sub>j</sub>. Flip-flop output Q transmits a control signal<o ostyle="single">end_transfert</o> at the logic gate ET2, allowing the clock signal h to pass to the shift2 shift input of the RDP processor register<sub>j</sub>. This control signal is also delivered to the management processor PG<sub>j</sub> to indicate the end of block transfer.
The operation of the assembly is as follows: the management processor PG<sub>j</sub>, after obtaining access to the central memory RAM via the BUSA bus, performs its block reading bi by providing the address of the block concerned and the read signal <o ostyle="single">r</o>. This signal triggers the activation of the TFRE transmission part<sub>j</sub> : the final edge of the read signal r causes the bi block to be loaded into the RDM memory register<sub>j</sub> by activating the signal <o ostyle="single">load1</o> and loading the value t + 1, corresponding to the size in bits of the block bi plus an additional bit called "start", in the down-counter register DC by the signal <o ostyle="single">load2</o> ; this has the effect of resetting the signal<o ostyle="single">borrow</o> and authorize the transfer clock H to supply, through the logic gate ET1 conditioned by this signal borrow, t + 1 clock pulses h: these pulses have the effect of shifting by the input shift1 t + 1 said from the RDM memory register<sub>j</sub> and to make reach the value 0 by the input down to the DC downcounter: the signal <o ostyle="single">borrow</o> is reset to zero and locks the operation of the TFRE transmission part<sub>j</sub>.
So the LS serial link<sub>j</sub>, initially in logical rest state 1, transmits bit 0 called start, then the t bits of block bi, and then returns to logical rest state 1, the last bit sent being the value l forced on the input RDM memory register series<sub>j.</sub>
Prior to the read request, the management processor PG<sub>j</sub> initialized the TFRR reception part '<sub>j</sub> by activating the signal <o ostyle="single">init</o> which has the effect of loading the RDP processor register<sub>j</sub> with t bits at 1 per input <o ostyle="single">load3</o> and to put the exit Q of rocker B in the logical state 1 by the entry <o ostyle="single">S</o>. This output Q then validates the logic gate ET2 which lets the clock signal h pass to the shift2 input of the RDP processor register.<sub>j</sub>. At each clock pulse this processor register provides a bit on its serial output which is memorized in flip-flop B. The first bit 0 which occurs has the effect of zeroing the output Q of flip-flop B and locking the signal clock h on gate ET2. This first bit 0 being the start bit which precedes the bi block, the latter is therefore trapped in the RDP processor register<sub>j</sub> when the management processor PG<sub>j</sub> is informed of the change of state of flip-flop B by the signal <o ostyle="single">end_transfert</o> : the PG management processor<sub>j</sub> just have to come and read this block bi on the parallel output of the RDP register<sub>j</sub>.
Writing a bi block to the central memory RAM requires the presence of a TFRE logic '<sub>j</sub>, identical to the TFRE logic<sub>j</sub>, associated with the RDP processor register<sub>j</sub>, and a TFRR logic<sub>j</sub>, identical to the TFRR 'logic<sub>j</sub>, associated with the RDM memory register<sub>j</sub>. In this case, the init signal of the TFRR logic<sub>j</sub> is connected to the write signal <o ostyle="single">w</o> : release of the RDM memory register<sub>j</sub> automatically reset the TFRR reception logic<sub>j</sub>.
This embodiment of the transfer control logic is only one possible example: the transmitter register can also be permanently shifted, and the receiver register activated for t clock pulses on detection of the start bit at the start. transfer.
The clock H can be connected to the two registers, or two independent local clocks can be used, synchronization being obtained conventionally by a so-called synchronization preamble.
The system shown in Figure 4 includes a dual memory shift register RDM1<sub>j</sub> and RDM2<sub>j</sub>, a split processor shift register RDP1<sub>j</sub> and RDP2<sub>j</sub>, two one-way serial links LS1<sub>j</sub> and LS2<sub>j</sub>, one connecting the RDM1 memory register<sub>j</sub> to the RDP1 processor register<sub>j</sub> so as to transmit the content from the first to the second, the other connecting the RDM2 memory register<sub>j</sub> to the RDP2 processor register<sub>j</sub> so as to transmit the content from the second to the first, and associated logic for controlling the transfer: TFRE1<sub>j</sub> for RDM1<sub>j</sub>, TFRR2<sub>j</sub> for RDM2<sub>j</sub>, TFRE2<sub>j</sub> for RDP2<sub>j</sub>, TFRR1<sub>j</sub> for RDP1<sub>j</sub>.
To read a block of information bi in central memory RAM, the management processor PG<sub>j</sub> initialized by signal <o ostyle="single">init</o> the TFRR1 logic<sub>j</sub> associated with the RDP1 processor register<sub>j</sub> then activates its request to read from RAM memory by the read signal <o ostyle="single">r</o>. This signal activates TFRE1 logic<sub>j</sub> associated with the RDM1 memory register<sub>j</sub> : this ensures the transfer on the link LS1<sub>j</sub> of the bi information block. The end of the transfer is detected by the TFRR1 logic<sub>j</sub> associated with the RDP1 processor register<sub>j</sub> which warns the management processor PG<sub>j</sub> of the arrival of the bi block by the signal <o ostyle="single">end_transfert</o>. The PG management processor<sub>j</sub> then transfers the content of the RDP1 processor register<sub>j</sub> in the MC cache<sub>j</sub>.
To write a bi memory block, the management processor PG<sub>j</sub> loads the RDP2 processor register<sub>j</sub> with the concerned bi block extracted from the MC cache memory<sub>j</sub>, which activates the transfer of this block on the link LS2<sub>j</sub>. The TFRR2 transfer logic<sub>j</sub> associated with the RDM2 memory register<sub>j</sub> ensures the good reception of this block. The PG management processor<sub>j</sub> is informed of the end of the transfer by the change of signal state <o ostyle="single">borrow</o> from TFRE2 transmission logic<sub>j</sub>. The PG management processor<sub>j</sub> then performs its write request which becomes effective upon activation of the write signal <o ostyle="single">w</o> ; this has the effect of transferring the content of the RDM2 register<sub>j</sub> in the main RAM memory and reset the TFRR2 logic for a next transfer<sub>j</sub>.
This device allows a simultaneous transfer of blocks in both directions and makes it possible to deal more quickly with bi block faults in the cache memory MC<sub>j</sub> when the latter is saturated; it also authorizes the establishment of a conventional block read anticipation mechanism.
In another embodiment presented in FIG. 5, the link LS<sub>j</sub> includes a single bi-directional link provided at each end with an LV logique and LV₂ validation logic constituted by a logic gate with 2 open-collector inputs OC₁ and OC₂, one of the inputs being connected to the serial output of the memory register RDM<sub>j</sub> for OC₁ gate and RDP processor register<sub>j</sub> for the gate OC₂, the other input being connected to the output Q of a control flip-flop BC₁ and BC₂; each of these is connected by its inputs<o ostyle="single">S</o> and <o ostyle="single">R</o> to the transfer logic TFR for the flip-flop BC₁ and TFR 'for the flip-flop BC₂.
Reads and writes are performed exclusively, on the sole initiative of the PG management processor<sub>j</sub>.
A memory read activates the read signal <o ostyle="single">r</o> which causes the BC₁ flip-flop to be set by its input <o ostyle="single">S</o>, reset being controlled, on the input <o ostyle="single">R</o>, by the TFR transfer logic at the end of the block transfer.
A memory write triggers an identical mechanism on the validation logic LV₂.
Other register / link combinations are possible, and in the case of a bi-directional link, it is possible in particular to use bi-directional shift registers receiving a transfer direction signal. This solution leads to the use of shift registers which are more complex in logic, therefore a priori less efficient in transfer speed.
The transfer speed must be very high, the RDM shift registers<sub>j</sub> and RDP<sub>j</sub>, their associated control logic TFR and TFR ', the validation logic LV₁ and LV₂, are chosen using fast technology (ECL, ASGA), and synchronized by a clock with frequency F at least equal to 100 MHz.
Another solution with multiplexed registers presented in FIG. 21 makes it possible, as will be understood later, to considerably reduce the quantity of efficient logic, therefore expensive, necessary.
The multiprocessor system of FIG. 1 was provided with both a common block address communication bus and serial data transfer links. FIG. 6 presents, as a variant, a multiprocessor system of the same general principle, but in which data and addresses pass through the serial links, in the absence of a common bus.
This system includes, in addition to the RDM memory registers<sub>j</sub>, ground-floor shift registers<sub>j</sub> able to memorize the addresses of the blocks requested and controlled by a TFR type logic<sub>j</sub>. In addition, an access management arbiter ABM is connected to the central memory RAM and to the complementary registers RDC<sub>j</sub> by their parallel output. Each TFR logic<sub>j</sub> is linked to this ABM arbiter of classical structure. The PG management processor<sub>j</sub> of each MC cache<sub>j</sub> is connected to a part of the parallel input of the RDP processor register<sub>j</sub>, in order to have access to it in writing.
To read a bi block in central memory RAM, the management processor PG<sub>j</sub> places the address of the requested block and the nature of the request (by a prefix bit: 1 = read, 0 = write) in the accessible part of the RDP processor register<sub>j</sub>, which has the effect of initiating the transfer of this information. The TFR transfer logic<sub>j</sub> detects the end of transfer on the RDC complementary register<sub>j</sub> and activates an operation request to the ABM arbitrator; it is responsible for serializing and processing block read requests in the central RAM memory by reading the address of the block requested in the RDC complementary register<sub>j</sub> corresponding to the transfer logic elected by the referee ABM, then by reading the block in central memory RAM which will then be loaded in the memory register RDM<sub>j</sub> and transmitted as before.
To write a block in central memory RAM, the management processor PG<sub>j</sub> connects the transmission of the address then of the block to be written through the RDP processor register<sub>j</sub>. The complementary DRC register<sub>j</sub> first receives the address and the nature of the request.
The TFR transfer logic<sub>j</sub> analyzes this request and validates the reception of the block in the RDM memory register<sub>j</sub> due to the nature of the request (writing). The TFR transfer logic<sub>j</sub> is notified of the end of transfer of the bi block and then transmits its service request to the ABM arbitrator. This request is processed, in turn, by said arbitrator who activates the writing of the block bi in memory.
Furthermore, the multiprocessor system shown in FIG. 7 includes means for managing shared data, making it possible to deal, statically, with the classic problem of maintaining the consistency of shared data. This system includes the resources of the system of FIG. 1 (same designations) with the following logics and additional resources:
A special BUSD word communication bus connects the CPUs<sub>j</sub> and the main memory RAM. LP partition logic<sub>j</sub> is associated with each CPU processor<sub>j</sub> ; each LP logic<sub>j</sub> is conventionally constituted by a set of register-comparator pairs connected in parallel on the adr bus address of the CPU processor<sub>j</sub>, in order to make a partition of the memory space of the central memory RAM into the area of non-shared data and of shared data, said logic LP<sub>j</sub> delivering for this purpose a signal p (indicating the nature of the data, shared or not). Decoding logic DEC is associated with the central memory RAM, itself arranged to be commanded in writing by word or by block by said logic DEC.
The decoding logic DEC is detailed in FIG. 8 and comprises a decoder DECL, receiving on its data input the word address part adrm of the address adr, and connected by its validation input to the output of a logic gate ET3 , each output i of said decoder being connected to a BFS output validation "buffer"<sub>i</sub>. The logic gate ET3 receives on its inputs the signal p and the signal<o ostyle="single">r</o> reversed. A DECE decoder is also connected by its data input to the adrm bus, and by its validation input at the output of a logic gate ET4, its outputs being connected to a set of logic gates OU1<sub>i</sub> in number equal to the number of words in a block. The logic gate ET4 receives on its inputs the signal p and the signal<o ostyle="single">w</o> reversed. The output of gate ET4 is also connected to a set of BFE₁, BFE input validation buffers<sub>i</sub>... The central memory RAM can be commanded in writing by word. Each word "slice" thus defined has its write command input w<sub>i</sub>. The output of each logic gate OU1<sub>i</sub> is connected to input w<sub>i</sub> of each word "slice" of the central memory RAM.
FIG. 8 also presents the detail of the addressing of the RDM memory registers<sub>j</sub>, which firstly includes a DECEB decoder connected by its data input to the common bus BUSA, with a view to receiving the number j from the processor concerned by the request from the central processing unit UC<sub>j</sub> ; this DECEB decoder is connected by its validation inputs to the output of an ET5 logic gate and by its outputs 1, 2 ... j to BV₁, BV validation "buffers"<sub>j</sub>... The logic gate ET5 receives on its inputs the signals p and <o ostyle="single">w</o> reversed. Similarly, a DECLB decoder is connected by its data input to field j of the common bus BUSA and by its validation input at the output of a logic gate ET6; the outputs 1, 2 ... j of this DECLB decoder are connected to the loading inputs ld₁, ld<sub>j</sub> RDM memory shift registers<sub>j</sub>. The logic gate ET6 receives on its inputs the signals p and<o ostyle="single">r</o> reversed.
The operation of the system is as follows: at each memory reference, the CPU processor<sub>j</sub> provides an address on its address bus address, and the nature of the request: reading <o ostyle="single">r</o> or writing <o ostyle="single">w</o>. It waits for data in the event of reading and provides data in the event of writing. The address adr crosses LP partition logic<sub>j</sub>, which indicates, by the signal p, whether the address adr belongs to an area of non-shared data (p = 0) or of shared data (p = 1). In the first case, the request is directed to the management processor PG<sub>j</sub> and is processed according to the operating mode described with reference to FIG. 1. In the second case, the request is directly routed to the common bus BUSA; the adr address bus includes additional address wires making it possible to identify the word concerned: the adr address consists of an adrb block address part and an adrm word address part. Thus, after agreement of the bus arbiter AB, the central memory RAM receives either a block transaction request (p = 0) and in this case, only the block part adrb of the address adr is significant, or a request for transaction word (p = 1) and, in this case, the entire adr address (adrb block and adrm word) is significant.
In case of block reading, p = 0 and r = 0, the logic gate ET6 validates the DECLB decoder which delivers a loading signal LD<sub>j</sub> on the RDM shift register<sub>j</sub>, allowing to load in this last the block read in central memory RAM at the address adrb by the read signal <o ostyle="single">r</o>.
In the case of block write, p = 0 and w = 0, the logic gate ET5 validates the DECEB decoder which delivers a validation signal to the BV buffer<sub>j</sub>, allowing the content of this register to be presented to the central memory RAM and thus to be written to the address adrb, the output of the logic gate ET5 providing the block write signal. The latter is broadcast on the write inputs w₁, w<sub>i</sub>, ... to the "slices" word of the central memory RAM through the logic gates OU1<sub>i</sub>.
When reading word, p = 1 and r = 0, the logic gate ET3 validates the DFCL decoder which delivers a validation signal to the BFS "buffer"<sub>i</sub>, allowing the requested word (of address adrm in the block adrb) whose reading is ensured by the signal <o ostyle="single">r</o>, to be directed to the special BUSD communication bus. This word is retrieved directly by the CPU processor<sub>j</sub> on its data entry.
In the case of word writing, p = 1 and w = 0, the logic gate ET4 validates the DECE decoder which provides at its output i a signal signaled through the logic gate OU1<sub>i</sub> to writing entry w<sub>i</sub> of the word "slice" of the central memory RAM concerned; this signal present at the input w<sub>i</sub> allows to write in this single "slice" word the word provided by the processor CPU<sub>j</sub> on the BUSD data bus. The content of this bus is presented in parallel on all the "slices" word of the central memory RAM, thanks to an activation of the "buffers" BFE<sub>i</sub> by the signal from the logic gate ET4.
An essential characteristic of the architecture of the invention is to present a minimum load of requests on the common bus BUSA. In the architecture shown diagrammatically in FIG. 7, the common bus BUSA is requested by block addresses and by word addresses. The frequency of requests for word addresses is a function of the rate of shared data and can lead to saturation of the common bus BUSA.
FIG. 9 alternatively presents a solution for reducing this load. The targeted system includes, in addition to the resources of FIG. 7, a BUSAM bus for word addresses, an arbitrator AB 'for arbitrating access conflicts to the BUSAM bus, an arbitrator ABM responsible for arbitrating access conflicts coming from the buses BUSA and BUSAM, and connected to a multiplexer MUX itself connected by its inputs to the two buses BUSA and BUSAM.
The operation of this system is as follows: as before, the LP partition logic<sub>j</sub> provides the signal p making it possible to identify the nature of the data handled.
If the request concerns non-shared data (p = 0), any information fault causes a block type memory request which passes through the common bus BUSA.
If the request concerns shared data (p = 1), the request is routed to the common BUSAM bus. Thus, the central memory RAM can receive simultaneous requests on the two buses BUSA and BUSAM, which must therefore be arbitrated. The arbitrator ABM allocates, in a conventional manner, access to the central memory RAM to one of the two requests and reconstructs the signal p from the origin of the request (p = 0 for BUSA, p = 1 for BUSAM). The signal p then controls, on the one hand, the multiplexer MUX which allows the signals of the bus concerned to pass through, on the other hand, the decoding logic DEC: we find ourselves in the situation of the previous system.
It will be noted that the load has moved from the common bus to the central memory RAM, since the request rate at the level of the latter remains the same and that its cycle time is of the same order of magnitude or even greater than that of the cycle. bus.
This solution is therefore only interesting if the central memory RAM consists of independent central memory banks organized according to the description given below with reference to FIG. 20: several transactions can in this case, if they affect different memory banks, take place simultaneously.
FIG. 10 presents a block diagram of an embodiment of architecture in accordance with the invention, in which the problem of shared data is dealt with dynamically. To this end, the device according to this embodiment comprises a spy bus processor PE<sub>j</sub>, coupled with a PGP parallel link management processor<sub>j</sub>. A PGS serial link management processor<sub>j</sub> is linked to the spy processor PE<sub>j</sub> by a FIFO queue<sub>j</sub>. A request management processor of the central processing unit PGU<sub>j</sub> is linked, on the one hand, to the processing processor CPU<sub>j</sub>, on the other hand, to the PGP parallel link management processors<sub>j</sub> and management of the PGS serial link<sub>j</sub>. The logic corresponding to the PG management processor<sub>j</sub> of each cache is in this embodiment exploded in the various processors presented above. Access to the MC cache memory<sub>j</sub> and to his RG repertoire<sub>j</sub> is regulated by a directory management processor and PGR cache<sub>j</sub>.
Finally, a management processor PGM of the central memory RAM is connected to the bus BUSA and to the central memory RAM, and to its shift registers RDM<sub>j</sub>.
The operation of the assembly is as follows: Each transaction on the common bus BUSA corresponds to a request to read or write bi block. PE bus spy processors<sub>j</sub> are activated by each data block read request. This operation carried out in the same cycle by all the spy processors will make it possible to guarantee the uniqueness of value of the shared data. The spy spy processor<sub>j</sub> has access to the RG directory<sub>j</sub>. The application function used for managing the MC cache memory<sub>j</sub> is in the described embodiment of the direct application type. Each element of the directory is a block descriptor which contains a "tag" field (block address), conventional bits of block states: a validation bit v and a modification bit m and two additional bits a to note that the block is known to the cache memory but still being transferred over the serial link, f to indicate that the block is in the FIFO queue and thus prevent it from being placed there several times.
The memory management processor PGM has, on the one hand, an AFIFO queue of bi block addresses and processor addresses, accessible in an associative manner, on the other hand, a directory of block state consisting of 2 bits per block ro and rw indicating the possible states of the following block:<ul id="ul0010" list-style="none"><li>. <maths id="math0002" num=""><math display="inline"><mrow><mtext>ro = rw = 0</mtext></mrow></math><img file="EP0346420B1_D0002.tif" /></maths> : block not yet broadcast,</li><li>. ro = 1; rw = 0: block already broadcast in reading: one or more copies of this block are in the cache memories,</li><li>. ro = 0; rw = 1: block broadcast in write mode: the updated copy of this block is in a cache memory.</li></ul>
The evolution of the block state bits is as follows, different according to the nature of the request from the processing processor CPU<sub>j</sub> : <ul id="ul0011" list-style="dash"><li>if the CPU processor<sub>j</sub> requests to read unshared data (program space or data explicitly not shared): the block is marked already broadcast in read mode (ro = 1; rw = 0) on the main memory side when transferring said block from the central RAM memory to the RDM register<sub>j</sub> and marked unmodified (m = 0) on the cache memory side in the same RG directory update cycle<sub>j</sub> cache memory (valid block). The spies did not react to the request on the common bus (the request having been made with the indication "data not shared").</li><li>if the CPU processor<sub>j</sub> makes a request to read data (a priori shared), the common bus BUSA is occupied for the time of the passage of address and request type information, the time of their processing by the PGM processor and the spies of the common bus PE<sub>j</sub>. In central RAM memory, this block can be:<ul id="ul0012" list-style="none"><li>1. Not yet released: <maths id="math0003" num=""><math display="inline"><mrow><mtext>ro = rw = 0</mtext></mrow></math><img file="EP0346420B1_D0003.tif" /></maths> . It is then transmitted to the central processing unit UC<sub>j</sub> and takes the unmodified state,</li><li>2. Already broadcast in reading: ro = 1; rw = 0. It is then transmitted to the central processing unit UC<sub>j</sub>. His condition does not change,</li><li>3. Already broadcast in writing: ro = 0; rw = 1. The updated copy of this block is in an MC cache memory<sub>i</sub>. The spy spy processor<sub>i</sub> associated with this cache noted the request from the central processing unit UC<sub>j</sub> when passing the address on the common bus and began transferring it to the central RAM memory as soon as possible on its LS serial link<sub>i</sub>. Pending its effective transfer, the memory management processor PGM places the request on hold in the associative queue which comprises a number of elements equal to the number of processors.</li></ul></li></ul>
During the request to read on the common bus BUSA, all the spy processors PE<sub>i</sub> reacted by consulting the RG directory<sub>i</sub> associated with their MC cache<sub>i</sub>. The BUSA common bus is only released when all the spy processors PE<sub>i</sub> had their access to the RG management directory<sub>i</sub>, which guarantees the same state of the block throughout the system. The processor which has the updated copy in its cache memory performs as soon as its serial link is free, the transfer of this block in the RDM register<sub>i</sub> and makes a block write request on the common bus which will have the effect of releasing the pending request in the AFIFO associative queue and of updating the block status bits.
Updating the block therefore only required writing to central RAM memory without activating the spy.<ul id="ul0013" list-style="dash"><li>if the CPU processor<sub>j</sub> requests the writing of a data in a block present in its cache-memory MC<sub>j</sub> with the status not modified, an informative write request must be sent on the common bus BUSA because it is possible that other cache memories MC<sub>i</sub> have this block with the same state. These other memories must be informed of the change of state. For this purpose, all spy processors PE<sub>i</sub> (activated by the informative writing broadcast on the common bus BUSA) consult their RG management directory<sub>i</sub> and invalidate this block, while the main memory notes at the same time the change of state of this block as well as the parallel management processor PGP<sub>j</sub>, in the RG management directory<sub>j</sub>. The release of the common bus BUSA by all the spy processors and the central memory RAM allows the processor CPU<sub>j</sub> write in its MC cache memory<sub>j</sub>, updating the status bit of the RG management directory<sub>j</sub> having been carried out.</li></ul>
If a central unit is waiting for access to the BUSA bus for the same request on the same block, its request is transformed into simple writing and then follows the block request protocol in writing.<ul id="ul0014" list-style="dash"><li>if the CPU processor<sub>j</sub> requests the writing of a data in a block absent from the cache memory MC<sub>j</sub>, this block is read in central memory RAM and brought into cache memory MC<sub>j</sub> so that the writing is made effective. In central RAM memory, this block can be:<ul id="ul0015" list-style="none"><li>1. Not yet released: <maths id="math0004" num=""><math display="inline"><mrow><mtext>ro = rw = o</mtext></mrow></math><img file="EP0346420B1_D0004.tif" /></maths> . The block is then sent on the LS serial link<sub>j</sub> to MC cache<sub>j</sub>. It takes the states ro = 0; rw = 1 in main memory and the modified state (m = 1) in the cache memory,</li><li>2. Already broadcast in reading: ro = 1; rw = 0. The block is sent on the LS serial link<sub>j</sub> to MC cache<sub>j</sub>. It takes the states ro = 0; rw = 1 in main memory and the modified state (m = 1) in the cache memory. When requesting on the common BUSA bus, the spy processors PE<sub>i</sub> noted the request and invalidated this block number in their MC cache<sub>i</sub>,</li><li>3. Already broadcast in writing: ro = 0; rw = 1. The request is put in the AFIFO associative queue and the common bus BUSA is released.</li></ul></li></ul>
The spy spy processor<sub>i</sub> MC memory<sub>i</sub>, holder of the updated copy, activates as soon as possible the transfer of the requested block from its cache memory MC<sub>i</sub> to central RAM memory. This block is then invalidated in the cache memory MC<sub>i</sub>.
CPU central unit<sub>j</sub> requests to write a block update in the following two cases:<ul id="ul0016" list-style="none"><li>a) the cache memory is full and purging a block requires updating this block in central memory,</li><li>b) a central unit CPU<sub>i</sub> is waiting for a block, the only updated copy of which is in the MC cache memory<sub>j</sub>. The spy spy processor<sub>j</sub> note the request and perform the purge of this block as soon as possible.</li></ul>
On the RAM central memory side, each update write request results in consultation of the AFIFO associative queue and, in the event of discovery of a central processing unit CPU<sub>i</sub> waiting for this block, loading this block in the RDM shift register<sub>i</sub> and updating the status bits corresponding to this block. This type of write request therefore does not request the spy processors.
The PGR directory management processor<sub>j</sub> which is added in this embodiment, allows the progression of the algorithm mentioned above, by coordinating the accesses to the directory of the cache-memory MC<sub>j</sub> which receives requests from three asynchronous functional units:<ul id="ul0017" list-style="none"><li>1. The CPU processing processor<sub>j</sub>, in order to read the instructions of the running program and to read or write the data handled by this program,</li><li>2. The common bus BUSA, in order to maintain the consistency of the data in the cache-memory MC<sub>j</sub>,</li><li>3. The LS serial link<sub>j</sub>, in order to load / unload a block of information from / to the central memory RAM.</li></ul>
Each of these requests accesses the RG management directory<sub>j</sub> from the cache. The serialization of these accesses on said management directory makes it possible to ensure the proper functioning of the aforementioned algorithm of information consistency in the cache memories. We thus obtain a strong coupling of requests at the level of the RG management directory.<sub>j</sub> but the synchronization which must exist at the level of the processing of these requests is sufficiently weak to envisage an asynchronous operation of the logic of processing of these requests, which leads to the following functional breakdown: The interface of each CPU<sub>j</sub> and its auxiliaries (PGS<sub>j</sub>, PGU<sub>j</sub>) with the common bus BUSA is composed of two parts having a mutually exclusive operation: the processor for managing the parallel link PGP<sub>j</sub> is responsible for requesting the common bus BUSA at the request of the request management processor PGU<sub>j</sub> or the PGS serial link management processor<sub>j</sub> and to control the common bus BUSA in writing. The PE bus spy processor<sub>j</sub> performs the spy function, which amounts to controlling the common BUSA bus in read mode. He frequently accesses the RG directory<sub>j</sub> MC memory<sub>j</sub>.
The PGS serial link management processor<sub>j</sub> manages the interface with the LS serial link<sub>j</sub>. It ensures the loading and unloading of blocks of bi information at the request of the request management processor PGU<sub>j</sub> and the PE bus spy processor<sub>j</sub>. It has infrequent access to the MC cache memory<sub>j</sub> and in the RG management directory<sub>j</sub> corresponding.
The PGU request management processor<sub>j</sub> tracks requests from the CPU<sub>j</sub>. It has very frequent access to the MC cache memory<sub>j</sub> and in the RG management directory<sub>j</sub>. This interface includes the possible logic of "MMU" ("Memory Management Unit") usually associated with the processing processor CPU<sub>j</sub>.
The PGR management processor<sub>j</sub> of the RG management directory<sub>j</sub> is the arbitrator responsible for allocating access to the MC cache memory<sub>j</sub>.
FIGS. 11, 12, 13, 14, 15, 16 and 17 represent, by way of example, embodiments of the various functional units of the device of FIG. 10. The designations of the signals or inputs and outputs of the units are chosen so usual. Signals of the same functionality which are generated in each functional unit from a basic signal will be designated by the same reference, for example: dnp = data not shared, dl = read request, maj = update, de = writing request, ei = informative writing. The device has several processors and the index j used until now aimed at a current processor and its auxiliaries; to simplify the description, this index has been omitted in these figures and it is understood that the description which follows relates to each of the functional units which are attached to each processing processor. Furthermore, the signals noted x_YZ define the name and the origin of the signal in the case where YZ = RG, MC, UC, and the source and the destination of the signal in the other cases, with Y and Z representing: U = PGU , R = PGR, P = PGP or PE, S = PGS.
The cache memory MC presented in FIG. 11 has for example a capacity of 16 KØ. It is organized into 16 fast RAM modules of 1 KØ MC₀, ... MC₁₅, each accessible on an edge of 4 bytes: the address bus of the cache-memory MC (denoted adr_MC) includes a block address part adr_bloc and a word address part in the adr_mot block. The address bus adr_MC is made up of 14 wires, making it possible to address the 16 KØs of the MC cache memory. The adr_bloc part has 8 wires for addressing the 256 block locations of the MC cache memory, and the adr_mot 6-wire part for addressing a word in the block, the size of which is 64 bytes in the example.
The address part adr_ block is connected to the address input of each of the memory modules MC₀ ... MC₁₅. The word address part adr_mot is connected to the input of two decoders DEC0 and DEC1 (only the 4 most significant bits of the address bus adr_mot are used: the address is a byte address and the cache has an access unit which is a 4 byte word). The read signal<o ostyle="single">r</o>_MC is delivered to each of the read inputs of the memory modules MC₀ ... MC₁₅ and to one of the inputs of a logic gate OU1. The other input of this logic gate OU1 receives the signal<o ostyle="single">block</o> reversed. The write signal<o ostyle="single">w</o>_MC is delivered to one of the two logic gate inputs OU2 and OU3. Logic gate OU2 receives the signal on its other input<o ostyle="single">block</o> reversed. Logic gate OU3 receives on its other input the signal<o ostyle="single">block</o>. The output of logic gate OU1 is connected to the validation input<o ostyle="single">in 1</o> of the decoder DEC1, and the output of rank i of this decoder DEC1 activates a "buffer" of validation BVL of rank i. The output of logic gate OU2 is connected to the validation input<o ostyle="single">in0</o> DEC0 decoder and BVE validation "buffers". The output of the logic gate OU3 is connected to logic gates ET1₀ ... ET1₁₅, which receive on their other input the output of corresponding rank from the decoder DEC0. The output i of each logic gate ET1₀ ... ET1₁₅ is connected to the write input<o ostyle="single">w₀</o>... <o ostyle="single">w₁₅</o> of each memory module MC₀ ... MC₁₅. A data bus connects each memory module MC₀ ... MC₁₅ to one of the BVL validation buffers and to one of the BVE validation buffers. The output of the RVL buffers and the input of the RVE buffers receive in parallel a datamot_MC data bus (connected to the PGU request management processor).
The operation of the example of the cache memory described above is as follows:
Case 1
The request comes from the PGS serial link management processor. This case is signaled by the presence of a logic zero state on the block signal.
In cache memory read, the PGS serial link management processor presents on the adr_MC address bus the address of the block location to be read (in this case, only the adr_bloc part of the adr_MC bus is used) and activates the signal reading <o ostyle="single">r</o>_MC. At the end of the access times, the block is available on the databloc_MC bus.
In write cache memory, the serial link management processor presents the address of the block to be written on the address bus adr_MC, on the databloc_MC data bus the data to be written there, and activates the line <o ostyle="single">w</o>_MC. Block signal zero state switches the signal<o ostyle="single">w</o>_MC to the write command inputs of the modules of the cache memory MC₀ ... MC₁₅, via the logic gates OU3 and ET1<sub>i</sub>. The information present on the databloc_MC data bus is written in cache memory at the end of the writing time.
Case 2
The request comes from the request management processor PGU of the processing processor CPU. This case is signaled by the presence of a logic state one on the block signal.
In cache-memory read, the management processor PGU presents the address of the requested word on the adr_MC bus, and activates the read signal <o ostyle="single">r</o>_MC. The block corresponding to the adr_bloc part is read in cache memory, and the requested word is routed, via one of the BVL validation buffers, to the datamot_MC data bus. The BVL validation buffer concerned is activated by the output of the decoder DEC1 corresponding to the word address adr_mot requested.
In cache memory write, the management processor PGU presents on the adr_MC bus the address of the word to be written, on the datamot_MC data bus the data to be written, and activates the write signal <o ostyle="single">w</o>_MC. The data present on the datamot_MC bus is broadcast on each cache-memory module, via the BVE buffers validated by the write signal. The write signal<o ostyle="single">w</o>_MC is then presented to the only memory module concerned. It is issued at the output of the DEC0 decoder corresponding to the address adr_mot concerned.
In the embodiment described above, the problems of access in byte and double byte, and of access in double byte and word straddling two memory modules are solved in the same way as in traditional computer systems and do not are not described here.
FIGS. 12a, 12b, 12c, 12d present, by way of example, the characteristics of a directory for managing the cache RG and an associated management processor PGR. FIG. 12a illustrates the logical structure of the address adr_RG in the hypothesis of an address space on 32 bits and with the characteristics of the cache memory described above. The -tag- field, component of the block address, is coded on 18 bits. The field -frame- is coded on 8 bits and makes it possible to address the 256 block locations of the cache memory MC. The last 6 bits define the word address in the block, in byte units.
Figure 12b shows the structure of the RG cache management directory, which is a simple fast RAM of 256 words of 22 bits. Each address word i contains the descriptor of the block registered in location i of the cache memory.
FIG. 12c schematizes the structure of the descriptor which comprises:<ul id="ul0018" list-style="dash"><li>an 18-bit tag field, defining the address of the block in the current location or block frame,</li><li>the validation bit v,</li><li>the modification bit m,</li><li>the end of transfer wait bit a,</li><li>the purge wait bit f.</li></ul>
FIG. 12d provides the structure of the PGR processor, which is none other than a conventional arbiter with fixed priority.
This arbitrator includes a LATCH register, three inputs of which receive signals respectively <o ostyle="single">rqst</o>_UR, <o ostyle="single">rqst</o>_PR, <o ostyle="single">rqst</o>_SR, also delivered respectively to logic gates ET2, ET3, ET4. The corresponding outputs of the LATCH register are connected to the inputs of a priority PRI encoder, the outputs of which are connected to the inputs of a DECPRI decoder. The outputs of rank corresponding to those of the LATCH register are connected to the signals<o ostyle="single">grnt</o>_UR, <o ostyle="single">grnt</o>_PR, <o ostyle="single">grnt</o>_SR as well as, inversely, respectively at the inputs of the logic gates ET2, ET3, ET4. The outputs of logic gates ET2, ET3, ET4 are connected to the inputs of logic gate NOU1. The output of the logic gate NOU1 is connected to a flip-flop B1, which receives on its input D the output<o ostyle="single">e0</o> of the PRI priority encoder. The entire device is synchronized by a general clock which delivers a signal h to one of the inputs clk of a logic gate ET5 and, inversely, to the clock input of the flip-flop B1. The exit<o ostyle="single">Q</o> of the flip-flop B1 is connected to the other input of the logic gate ET5. The output of logic gate ET5 is connected to the load input of the LATCH register.
The functioning of this arbitrator is as follows: in the absence of any request on the lines <o ostyle="single">rqst</o>, flip-flop B1 permanently stores the state of the line <o ostyle="single">e0</o>, inactive, and thus validates through logic gate ET5 the loading of the LATCH register.
The arrival of a signal <o ostyle="single">rqst</o> causes the clock to lock and activate the signal <o ostyle="single">grnt</o> associated with signal <o ostyle="single">rqst</o>, until the latter is deactivated: the arbitrator is frozen in his state for the duration of the current transaction.
The PGU request management processor, represented in FIG. 13, constitutes an interface between the processing processor CPU and:<ul id="ul0019" list-style="dash"><li>on the one hand, the various processors with which it must exchange information: PGP parallel management processor, PGS serial management processor, PGR cache directory management processor,</li><li>on the other hand, the management directory of the cache memory RG and the cache memory MC.</li></ul>
The processing processor CPU triggers the activity of the request management processor PGU by activating the signal <o ostyle="single">ace</o> ("address strobe"). This signal validates the address bus adr_CPU, the read signals<o ostyle="single">r</o>_CPU and writing <o ostyle="single">w</o>_CPU as well as the fc_CPU function lines of the processing processor CPU. The processing processor CPU then waits until the request is acknowledged by the signal<o ostyle="single">dtack</o>_CPU.
The signal <o ostyle="single">ace</o> is connected to the input of a differentiating circuit D10. The output of this circuit is connected to one of the three inputs of an ET12 logic gate, the other two inputs respectively receiving signals<o ostyle="single">ack</o>_US and <o ostyle="single">ack</o>_UP. This last signal is also issued at the entrance<o ostyle="single">R</o> a B13 scale. Entrance<o ostyle="single">S</o> of the flip-flop B13 receives the output of a logic gate NET10. The output of logic gate ET12 is connected to the input<o ostyle="single">S</o> a B11 scale at the entrance <o ostyle="single">S</o> of scale B10 and at the entrance <o ostyle="single">clear</o>10 an SR10 shift register. The exit<o ostyle="single">Q</o> of flip-flop B11 provides the signal <o ostyle="single">rqst</o>_UR. The B11 scale receives on its input<o ostyle="single">R</o> phase ϑ13 reversed and flip-flop B10 phase ϑ11 reversed. The output Q of the flip-flop B10 is connected to the serial input serial_in10 of the register SR10. The shift register SR10 receives on its clk10 input the clock signal -h- and on its validation input<o ostyle="single">in10</o> the signal <o ostyle="single">grnt</o>_UR.
Signal activation <o ostyle="single">ace</o> triggers the operation of the D10 differentiation circuit. The impulse produced by this circuit crosses the logic gate ET12, puts in the logic state one flip-flops B10 and B11 by their input<o ostyle="single">S</o>, and also performs the reset of the shift register SR10 by its input <o ostyle="single">clear</o>10.
The flip-flop B10 and the shift register SR10 constitute the logical subset "phase distributor" DP_U. The activation of this phase distributor is triggered by setting the flip-flop B10 to one and resetting the shift register SR10. If the shift register is validated by the presence of a zero level on its input<o ostyle="single">in10</o>, then the next clock pulse h on the input clk of the shift register produces the shift of one step of said register.
The logic state one of the flip-flop B10 is reflected on the output ϑ11 of the shift register SR10 by its serial input serialin10. The output ϑ11, called phase ϑ11, inverted, resets flip-flop B10 to zero by its input<o ostyle="single">R</o>. Thus, a single bit is introduced into the shift register SR10 each time the phase distributor DP_U is activated. At each clock pulse h, this bit will shift in the shift register SR10 and produce the consecutive disjoint phases ϑ11, ϑ12, ϑ13.
Putting flip-flop B11 in logic state causes the signal to be activated <o ostyle="single">rqst</o>_UR. This signal is sent to the PGR directory management processor. The latter, as soon as possible, will grant access to the RG management directory and to the MC cache memory by activating the signal<o ostyle="single">grnt</o>_UR, which will validate all the buffer buffers BV10, BV11 and BV12 located respectively on the buses of the management directory and on the buses of the cache memory. This signal<o ostyle="single">grnt</o>_UR also validates the phase distributor which will therefore produce the phases ϑ11, ϑ12, ϑ13 sequentially.
Phase ϑ11 corresponds to a time delay making it possible to read the descriptor of the block requested by the processing processor CPU in the management directory RG, addressed by the frame field of the address adr_CPU and connected to the bus adr_RG through buffers of passage BV10 . The signal<o ostyle="single">r</o>_RG is always active at the input of a BV10 buffer, the signal w_RG always inactive at the input of a BV10 buffer. At the end of the time delay, the descriptor is returned to the PGU processor via the data_RG bus. The tag part of this descriptor and the validation bit v are delivered to one of the comparison inputs of a comparator COMP10, the other input being connected to the tag part of the address adr_CPU. The bit next to the validation bit v is always one. The comparator COMP10 is permanently validated by the presence of a level one on its input en11.
Access time to the management directory RG and clock frequency h are so related that at the end of phase ϑ11, the output eg10 of comparator COMP10 is positioned and provides the information "the requested block is present in cache or not in cache ".
If the requested block is present in the cache memory MC (eg10 = 1) then the signal eg10, delivered to one of the two inputs of the logic gate ET10, provides a signal calibrated by phase ϑ12, present on the other input of logic gate ET10.
This calibrated signal present on the output of logic gate ET10, is linked to the inputs of logic doors NET10, NET11, NET12.
The logic gate NET10 receives on its inputs, in addition to the output of the logic gate ET10, the inverted state bit m from the descriptor, and the inverted write request signal <o ostyle="single">w</o>_CPU.
The activation of the NET10 logic gate corresponds to the state "request to write a word in a block present in the cache and which has not yet been modified (m = 0)". The logic gate NET10 output is connected to the input<o ostyle="single">S</o> a B13 scale. Activation of logic gate NET10 puts flip-flop B13 in logic state one, which triggers an informative write request by the line<o ostyle="single">rqst</o>_UP to the PGP parallel link management processor. The address of the block concerned is provided by the lines adr_bloc_UP, derived from the lines adr_CPU.
The PGU request management processor has completed the first part of its task: the RG management directory and the cache memory are freed by deactivation of the signal <o ostyle="single">rqst</o>_UR, consequence of the arrival of the reversed phase ée13 on the input <o ostyle="single">R</o> of scale B11.
The mechanism of informative writing is described in the paragraph "PGP parallel link management processor", and has the effect of putting the requested block in the modified (m = 1) or invalid (v = 0) state. It will be noted that the release of the management directory RG and the cache memory MC by the request management processor PGU is necessary so that the parallel link management processor PGP can have access to it. The end of the "informative write" operation is signaled to the PGU request management processor by the activation of the signal <o ostyle="single">ack</o>_UP, which has the effect of resetting the flip-flop B13 and activating, through the logic gate ET12, the flip-flop B11 and the phase distributor: the cycle initially started by the signal <o ostyle="single">ace</o> is repeated, but the sequence resulting from the activation of gate NET10 will not repeat a second time in this request cycle.
The logic gate NET11 receives on its inputs, in addition to the output of the logic gate ET10, the status bit m from the descriptor, and the inverted write request signal <o ostyle="single">w</o>_CPU.
The activation of the NET11 corr gate waits for the state "request to write in a block present in the cache and which has already been modified".
The output of gate NET11 is connected, through one of the BV11 buffers, to the write signal <o ostyle="single">w</o>_MC of the MC cache memory. This signal makes it possible to write into the cache memory, at the address present on the adr_MC bus, connected to the adr_CPU bus, via the BV11 buffers, the data present on the data_MC bus, connected to the data_CPU bus via the BV12 bidirectional buffers. The direction of activation of these buffers is provided by the signal<o ostyle="single">w</o>_MC.
The output of gate NET11 is also connected to one of the inputs of logic gate ET11, which thus sends the signal <o ostyle="single">dtack</o>_CPU to the CPU processing processor. The write operation in the cache is done in parallel with the activation of the signal<o ostyle="single">dtack</o>_CPU, which conforms to the usual specifications of processing processors.
The operation ends with the release of the RG management directory and the MC cache memory by deactivating the signal <o ostyle="single">rqst</o>_UR, consequence of the arrival of the reversed phase invers13 on the flip-flop B11.
The logic gate NET12 receives on its inputs, in addition to the output of the logic gate ET10, the inverted read request signal <o ostyle="single">r</o>_CPU. The activation of the NET12 logic gate corresponds to the state "request to read a word in a block present in the cache".
The sequence of operations is identical to the previous operation, with the only difference of the activated signal (<o ostyle="single">r</o>_MC rather than <o ostyle="single">w</o>_MC) associated with the direction of data transit on the data_CPU and data_MC buses.
If the requested block is absent from the cache memory (eg10 = 0), then the signal eg10, inverted, connected to one of the two inputs of the logic gate NET13, provides a signal calibrated by phase ϑ12, present on l another input of the NET13 logic gate. The output of the NET13 logic gate is connected to the input<o ostyle="single">S</o> of scale B12. This calibrated signal forces flip-flop B12 to one, which has the effect of issuing a service request<o ostyle="single">rqst_US</o> to the PGS serial link management processor. This processor also receives the address of the block to be requested on the lines adr_bloc_US and the nature of the request on the lines<o ostyle="single">w</o>_US, <o ostyle="single">r</o>_US and fc_US.
The PGU request management processor has completed the first part of its task: the RG management directory and the MC cache memory are freed by deactivation of the line <o ostyle="single">rqst</o>_UR, consequence of the arrival of the reversed phase invers13 on the flip-flop B11.
The mechanism for updating the cache is described in the paragraph "PGS serial link management processor".
It will be noted that the release of the management directory RG and of the cache memory MC is necessary so that the management processor of the serial link PGS can have access to it.
The update of the cache is signaled to the request processor PGU by the activation of the signal <o ostyle="single">ack</o>_US. This signal is issued at the entrance<o ostyle="single">R</o> from scale B12 and at the entrance to gate ET12. It thus has the effect of resetting the flip-flop B12 and activating through the logic gate ET12, the flip-flop B11 and the phase distributor: the cycle initially started by the signal<o ostyle="single">ace</o> recurs, but this time successfully due to the presence of the block in the cache memory.
The serial management processor PGS shown by way of example in FIG. 14 is responsible for managing the serial link LS and for this reason carrying out block transfer requests between central memory RAM and cache memory MC and carrying out updates corresponding updates in the RG management directory. It processes in priority the requests coming from the PE spy processor waiting in a FIFO queue. It also processes requests from the PGU request processor.
This PGS serial link management processor comprises a flip-flop B20 which receives the signal on its data input D <o ostyle="single">empty</o> from the FIFO queue and on its clock input the output <o ostyle="single">Q</o> of a scale B22. A flip-flop B21 receives on its data input the output of a logic gate OU20. This OU20 logic gate validates the signal<o ostyle="single">rqst</o>_US, received at one of its two inputs, the other input being connected to the signal <o ostyle="single">empty</o>. The clock input of flip-flop B21 comes from output Q of flip-flop B22. The exit<o ostyle="single">Q</o> of the flip-flop B22 is looped back to its data input D, which conditions it as a divider by two. The clock input of the flip-flop B22 is connected to the output of a logic gate ET20, which receives on one of its inputs the general operating clock h and on the other input a validation signal. This validation signal comes from the output of an ET24 logic gate, which receives the outputs on its two inputs respectively<o ostyle="single">Q</o> and Q of flip-flops B20 and B21.
The B20 scale receives on its input <o ostyle="single">R</o> phase ϑ25, inverted, coming from a DP_S phase distributor. The flip-flop B21 receives on its input<o ostyle="single">S</o> the output of a NOU20 logic gate. This NOU20 logic gate receives on its two inputs the respective outputs of logic gates ET22 and ET23. The logic gate ET22 receives on its inputs the phase ϑ25 from the phase distributor and the major signal from a logic gate ET29. The logic gate ET23 receives on its inputs the phase ϑ27 from the phase distributor and the inverted shift signal.
The set of circuits B20, B21, B22, ET20, ET22, ET23, OU20, ET24, NOU20 constitutes a fixed priority arbiter ARB_S. Its operation is as follows: the flip-flop B22 provides on its outputs<o ostyle="single">Q</o> and Q of the alternating signals of frequency half that of the general clock. These signals validate alternately the flip-flops B20 and B21. If a service request is present on one of the inputs of flip-flops B20 or B21, then these alternating signals store the request in the corresponding flip-flop (B20 for the spy processor PE, B21 for the request management processor PGU) which, in turn, locks alternating operation. Note that the signal<o ostyle="single">rqst</o>_US coming from the request management processor PGU is conditioned by the signal <o ostyle="single">empty</o> through the OU20 gate: this signal is thus only taken into account if the FIFO queue is empty. The flip-flop B20 (respectively B21) is not reset until the transaction to be carried out is completed.
The sampling of a request on one or other of the flip-flops B20 and B21 results in a change of state of the output of the logic gate ET24. The output of the logic gate ET24 is also connected to a differentiator circuit D20 which delivers a pulse when the state of the output of the logic gate ET24 changes. The output of the differentiator circuit is connected on the one hand to the phase distributor DP_S (input<o ostyle="single">S</o> a B36 scale and <o ostyle="single">clr</o>20 of a shift register SR20) of the serial link management processor and on the other hand to one of the two inputs of the logic gate OU22. The output of the logic gate is connected to the inputs<o ostyle="single">S</o> of two scales B24 and B25. The B24 scale receives on its input<o ostyle="single">R</o> the output of a logic gate NOU21 and the flip-flop B25 the signal <o ostyle="single">grnt</o>_SR. The logic gate NOU21 receives on its two inputs the outputs of logic doors ET36 and ET37. The logic gate E36 receives on its inputs the phase ϑ23 coming from the phase distributor and the maj signal, and the logic gate E37 receives the phase ϑ27 coming from the phase distributor and the inverted maj signal.
The pulse from the differentiator circuit D20 thus initializes the phase distributor and puts one flip-flops B24 and B25 into logic state through logic gate OU22.
The DP_S phase distributor consists of the shift register SR20 and the flip-flop B36. Its operation is identical to that described in the paragraph concerning the PGU request management processor.
The exit <o ostyle="single">Q</o> of flip-flop B24 is connected to the signal <o ostyle="single">rqst</o>_SR, for the PGR directory processor. Its activation triggers a service request to this processor, which responds by line<o ostyle="single">grnt</o>_SR, connected to the input <o ostyle="single">R</o> of scale B25. The Q output of flip-flop B25 is connected to one of the inputs of a logic gate OU23. The output of logic gate OU23 is connected to the input<o ostyle="single">in20</o> of the shift register SR20.
The logic set B24 and B25 constitutes a resynchronization logic RESYNC_S between asynchronous units. Its operation is as follows:
A service request <o ostyle="single">rqst</o>_SR to the PGR directory management processor is done by activating flip-flops B24 and B25 through logic gate OU22, which authorizes two origins of activations. The logic specific to the PGR directory management processor ensures a response<o ostyle="single">grnt</o>_SR in an indefinite time, which resets the flip-flop B25. The B24 scale maintains its request until it is reset to zero by activating its input<o ostyle="single">R</o>. In return, the PGR directory management processor deactivates its line<o ostyle="single">grnt</o>_SR: the resynchronization logic is ready to operate for a next service request. The output Q of flip-flop B25 is used to block the phase distributor by acting on its input<o ostyle="single">in20</o> via logic gate OU23: resetting the flip-flop B25 releases the phase distributor which will supply the first phase ϑ21 during the next active transition of the general clock h, connected to the clk20 input of the phase distributor.
Line <o ostyle="single">grnt</o>_SR is also linked to the validation buffers BV20, BV25 and BV21, BV22 which open access to the RG management directory and to the cache memory MC respectively.
If the flip-flop B20 is active, then the current transaction is a block purge requested by the spy processor PE via the queue FIFO. The exit<o ostyle="single">Q</o> of the B20 scale is connected to BV24 validation buffers. These buffers receive on their input the output of a REG20 register. The exit<o ostyle="single">Q</o> of the flip-flop B20 is connected to one of the two inputs of a logic gate OU21, which receives on its other input the output of the differentiator circuit D20. The output of logic gate OU21 is connected to the input<o ostyle="single">load20</o> REG20 register and at the entrance <o ostyle="single">read20</o> from the FIFO queue.
Thus, the activation of the B20 flip-flop causes:<ul id="ul0020" list-style="none"><li>1. The initialization of the phase distributor,</li><li>2. A request for access to the RG management directory and to the MC cache memory,</li><li>3. The loading of the element at the head of the FIFO queue in the REG20 register and the advance of the queue,</li><li>4. The validation of the BV24 buffer: the adr_X bus contains the address of the block to be purged. The nature of the operation (update) will be found from the combination of the bits v and m (major signal).</li></ul>
If the flip-flop B21 is active, then the transaction in progress comes from an information fault in the cache memory MC. The exit<o ostyle="single">Q</o> of flip-flop B21 is connected to validation buffers BV23. These buffers receive information on their input from the PGU request management processor.
Thus the activation of the flip-flop B21 causes:<ul id="ul0021" list-style="none"><li>1. The initialization of the phase distributor,</li><li>2. A request for access to the RG management directory and to the MC cache memory,</li><li>3. The validation of the BV23 buffers: the adr_X bus contains the address of the block which caused the information fault in the MC cache memory, and the nature of the request: read or write, shared or unshared data (lines fc_US) .</li></ul>
The "frame" field of the adr_X bus is linked to the adr_RG address lines of the RG management directory through BV20 validation buffers. The exits<o ostyle="single">Q</o> scales B26 and B27 are respectively connected to the lines <o ostyle="single">r</o>_RG and <o ostyle="single">w</o>_RG of the management directory through the BV20 validation buffers. The B26 scale receives on its input<o ostyle="single">S</o> the output of logic gate OU22 and on its input <o ostyle="single">R</o> reverse phase ϑ22 from the phase distributor. The B27 scale receives on its input<o ostyle="single">S</o> the output of a NOU22 logic gate and on its input <o ostyle="single">R</o> the output of a NOU23 logic gate. The logic gate NOU22 receives on its two inputs the respective outputs of logic gates ET25 and ET26, themselves receiving on their inputs, for the gate ET25 the phase ϑ22 and the maj signal and for the gate ET26 the phase ϑ26 and the maj signal . The logic gate NOU23 receives on its two inputs the respective outputs of logic gates ET27 and ET28, themselves receiving on their inputs, for the gate ET27 the phase ϑ23 and the shift signal inverted and for the gate ET28 the phase ϑ27 and the signal reverse shift.
The tag field of the adr_X bus is linked to the tag field of the data_RG lines through BV25 validation buffers. These receive on their validation input the output of an OU24 logic gate which receives on its inputs the signal<o ostyle="single">grnt</o>_SR and the exit <o ostyle="single">Q</o> of scale B27.
The inputs D of flip-flops B28 and B29 are respectively connected to the lines of validation bit v and modification bit m of the data_RG bus. The clock inputs of these flip-flops B28 and B29 are connected to phase ϑ22 of the phase distributor. The Q output of flip-flop B28 and the Q output of flip-flop B29 are connected to the two inputs of an ET29 logic gate, which provides the maj signal on its output. An OU25 logic gate receives the maj signal and the output at its inputs<o ostyle="single">Q</o> a B30 scale. The output of the logic gate OU25 is connected to the selection input sel20 of a MUX20 multiplexer, which receives on its two data inputs the output Q of the flip-flop B28 (bit v) and the constant 0, the constant zero being chosen when the sel20 selection input is in logic state one. The output of the MUX20 multiplexer is connected to the validation bit line v of the adr_X bus. The wait bit line -a- of the adr_X bus is forced to logic zero. The said modification -m- is linked to the reading line <o ostyle="single">r</o>_adr_x of the adr_X bus.
The logic assembly described above constitutes the logic for accessing and updating the RG management directory. Its operation is as follows: activation of the signal<o ostyle="single">grnt</o>_SR, authorizing access to the management directory and the cache memory, validates the BV20 validation buffers. Reading of the descriptor concerned is commanded from the start of the access authorization until the arrival of the phase ϑ22, moment of storage of the bits -v- and -m- in flip-flops B28 and B29. The combined state of these two bits produces, through logic gate ET29, the maj signal which conditions the continuation of the operations.<u style="single">1st case</u> : maj = 1. This case occurs when the validation bit is equal to 1 and the modification bit is equal to 1. This case corresponds either to a request to purge the spy processor PE block, or to a fault of information found by the PGU request management processor on an occupied and modified block location: in both cases, the block concerned must be written in central memory RAM.
To this end, the field -frame- of the adr_X bus is linked to the adr_MC lines via the BV21 validation buffers. The exits<o ostyle="single">Q</o> flip-flops B30 and B31 are respectively connected to the lines <o ostyle="single">r</o>_MC and <o ostyle="single">w</o>_MC from the MC cache memory, through the BV21 validation buffers. Line<o ostyle="single">block</o> is forced to zero via one of the BV21 validation buffers. The data_MC data lines of the cache memory are connected, via the bi-directional validation buffers BV22, to the inputs of the shift register RDP and to the outputs of the validation buffers BV26, which receive on their inputs the output lines of the RDP shift register. BV22 buffers are validated by the line<o ostyle="single">grnt</o>_SR, and their direction of validation controlled by the output <o ostyle="single">Q</o> of scale B31. The B30 scale receives on its inputs<o ostyle="single">S</o> and <o ostyle="single">R</o> respectively the phases ϑ21 and ϑ23, inverted, coming from the phase distributor. The B31 scale receives on its inputs<o ostyle="single">S</o> and <o ostyle="single">R</o> the logic gate outputs ET32 and ET33 respectively. The logic gate ET32 receives on its inputs the phase ϑ26 and the shift signal inverted, the logic gate ET33 the phase ϑ27 and the shift signal inverted. A NET20 logic gate receives on its inputs the phase maj23 and the major signal. The output of an ET35 logic gate controls the BV26 validation buffers, and receives on its two inputs respectively the inverted shift signal and the Q output of the flip-flop B31. The output of the logic gate NET20 is connected to the signal <o ostyle="single">load21</o> RDP shift register and input <o ostyle="single">S</o> a B32 scale. Entrance<o ostyle="single">R</o> of this flip-flop B32 receives the signal <o ostyle="single">fin_transfert_maj</o> from the TFR logic associated with the RDP shift register. The Q output of flip-flop B32 is connected to one of the inputs of logic gate OU23.
The logic described above makes it possible to purge a block from the cache memory. Its operation is as follows: In parallel with access to the RG management directory, a reading of the MC cache memory is activated by the line <o ostyle="single">r</o>_MC, from flip-flop B30, during phases ϑ21 and ϑ22. At the end of this reading, the data read, representing the block to be unloaded, are present at the input of the shift register RDM. Activation of the maj signal, the state of which is known at the start of phase ϑ22, causes:<ul id="ul0022" list-style="none"><li>1. Invalidation of the block in the management directory: the sel20 input of the MUX20 multiplexer being in logic state one, the value zero is forced on the validation bit, the descriptor being written in the RG management directory with the signal activation <o ostyle="single">w</o>_RG, controlled by flip-flop B27 during the cycle ϑ22,</li><li>2. The loading of the shift register RDM and the activation of the transfer, during the transition from phase ϑ22 to phase ϑ23,</li><li>3. The setting to logic one of flip-flop B32, which will block the phase distributor on phase ϑ23 until the end of the transfer, signaled by the signal <o ostyle="single">fin_transfert_maj</o> based on TFR logic,</li><li>4. The release of the RG management directory and the MC cache memory by resetting the flip-flop B24 to phase ϑ23.</li></ul>
Thus, access to the directory is freed (therefore accessible for the spy processor PE), the update transfer is in progress, and the phase distributor blocked on ϑ23.
As soon as the transfer is complete, the block is waiting in the shift register RDM and it is necessary to activate the processor for managing the parallel link for the writing in central memory RAM to be effective.
For this purpose, the flip-flop B33 is linked by its output Q to the service request line <o ostyle="single">rqst</o>_SP to the processor for managing the PGP parallel link. The output Q is connected to one of the inputs of the logic gate OU23, the input<o ostyle="single">S</o> at phase ϑ24 reversed from the phase distributor and the input <o ostyle="single">R</o> at the signal <o ostyle="single">ack</o>_SP. The adr_X bus is connected to the adr_bloc_SP bus of the PGP parallel link management processor. One of the lines of the adr_bloc_SP bus receives the maj signal in order to indicate the nature of the request: update.
As soon as the phase distributor is released by the signal <o ostyle="single">fin_transfert_maj</o>, the next active transition from the general clock h causes the transition from phase ϑ23 to phase ϑ24. Phase ϑ24 causes a service request to the processor for managing the PGP parallel link (activation of the line<o ostyle="single">rqst</o>_SP) and the blocking of the phase distributor until the signal is acknowledged <o ostyle="single">ack</o>_SP. At this time, the writing in update will have been effectively carried out by the processor for managing the parallel link PGP. The phase distributor, on the next active transition of clock h, will go from phase ϑ24 to phase ϑ25.
The update of the central memory RAM is finished: the flip-flop B20 is set to zero by activation of its input <o ostyle="single">R</o> by the reversed ϑ25 phase. Flip-flop B21 is set to one by activating its input<o ostyle="single">S</o> by phase ϑ25, conditioned by the major signal by means of logic gate ET22 via logic gate NOU22. In the event of a lack of information in the cache memory, purging a block is only the first part of the request: the request<o ostyle="single">rqst</o>_US is still present, but the release of the flip-flop B21 will make it possible to take into account any requests for updates pending in the FIFO queue. As soon as the FIFO queue is empty (empty = 0), the whole cycle described above is repeated, this time with the validation bit set to zero. We then find ourselves in the following case.<u style="single">2nd case</u> : maj = 0. This case occurs when the validation bit is equal to zero (perhaps as a result of a block purge) or when the validation bit is equal to one but the modification bit is equal at zero: the updated copy of this block is already in central RAM memory.
So the service request <o ostyle="single">rqst</o>_SR will in turn lead to an access agreement to the management directory RG and to the cache memory MC with reading of the descriptor, memorization of bits m and v, generation of the maj signal, rewriting of the descriptor with v = 0 (maj is at logic state zero, flip-flop B31 is in logic state zero and its output Q therefore equals one, which imposes a logic signal at state one on the input sel20 of the multiplexer MUX20 and therefore forces the constant 0) and release of access to the management directory RG and of the cache memory MC. These operations are effective upon activation of phase ϑ23.
The requested block should now be read from the central RAM memory. For this purpose, the inverted shift signal is received at one of the inputs of a logic gate NET21, which receives on its other input the phase ϑ25. The output of the NET21 logic gate is connected to the inputs<o ostyle="single">S</o> flip-flops B34 and B35. The B34 scale receives on its input<o ostyle="single">R</o> the signal <o ostyle="single">end_reception</o> from the TFR transfer logic. Its output Q is connected to the differentiator circuit D21. This circuit is connected to logic gate OU22. Entrance<o ostyle="single">R</o> of scale B35 is connected to the signal <o ostyle="single">grnt</o>_SR and the Q output at one of the inputs of the logic gate OU23.
The reading of a block in central memory RAM and its transfer in the cache memory are carried out as follows: the transition from phase ϑ23 to phase ϑ24 triggers a service request to the processor for managing the parallel link by activating line <o ostyle="single">rqst</o>_SP, from rocker B33. The type of operation is this time a read (r_adr_X = 0) or a write (w_adr_X = 0) and the adr_X bus provides the address of the block requested on the adr_bloc_SP bus. The phase distributor is blocked until the acknowledgment signal arrives<o ostyle="single">ack</o>_SP: the read or write request was made to the central memory RAM by the processor for managing the parallel link PGP, and the block was at the same time validated and marked "pending" by this same processor. The transfer is therefore in progress, from the central memory RAM to the shift register RDP.
The released phase distributor then supplies phase ϑ25. This phase, via the logic gate NET21 which receives the inverted shift signal at its other input, will put the flip-flops B34 and B35 into logic state. The B35 scale blocks the phase distributor. The flip-flop B34 is reset to zero upon arrival of the block requested in the RDP register, signaled by the activation of the line<o ostyle="single">end_reception</o> resulting from the TFR transfer logic, which triggers the activation of the differentiator circuit D21 and a request for access to the management directory RG and to the cache memory MC, by activation of the line <o ostyle="single">rqst_SR</o> from rocker B24.
The access agreement, signaled by the line <o ostyle="single">grnt</o>_SR, frees the phase distributor by resetting the flip-flops B25 and B35 and opens access to the RG management directory and to the MC cache memory. The phase distributor, on the next active transition of the general clock h, supplies the phase ϑ26.
B27 flip-flop, connected to the signal <o ostyle="single">w</o>_RG of the management directory, is active from ϑ26 to ϑ27, which allows updating the management directory with:<ul id="ul0023" list-style="dash"><li>date_RG bus tag field = adr_X bus tag field,</li><li>validity bit v = 1 (shift = 0 and the output Q of the flip-flop B31 is at zero: the multiplexer MUX20 lets v pass, which was forced to one by the processor for managing the parallel link PGP, or already reset to zero by the spy processor PE),</li><li>modification bit m = line status <o ostyle="single">r</o>_adr_X of the adr_X bus (write request causes m = 1, read request m = 0),</li><li>transfer wait bit a = 0,</li><li>bit f = 0.</li></ul>
Flip-flop B31 activates the signal <o ostyle="single">w</o>_MC of the MC cache memory from ϑ26 to ϑ27, which allows the contents of the RDM shift register to be written (the BV26 buffers are validated by the output <o ostyle="single">Q</o> of the flip-flop B31 and the inverted shift signal) in the correct location (field -frame- of the bus adr_X connected to the bus adr_MC) of the cache memory, lives the validation buffers BV22, controlled in the direction of transfer by the flip-flop B31.
When phase ϑ27 arrives, the update of the management directory and the cache memory is complete. The arrival of phase provoque27 causes the reset of flip-flop B24, which frees access to the management directory RG and to the cache memory MC and activation of the output of the logic gate ET23, resetting a flip-flop B21 and activating the signal<o ostyle="single">ack</o>_US to the PGU request management processor: the request is finished.
Furthermore, the spy processor PE is responsible for maintaining the consistency of the data in the cache memory MC; by way of example, FIG. 15 shows an architecture of this processor. This is triggered at each active transition of the signal<o ostyle="single">valid</o> from the BUSA common bus.
If the signal <o ostyle="single">valid</o> is activated by a processor for managing the parallel link PGP other than that associated with the spy processor PE, then the task of the latter is as follows, depending on the type of request:<ul id="ul0024" list-style="dash"><li>non-shared data block read request or block update write request: none,</li><li>shared data block read request:<ul id="ul0025" list-style="none"><li>. absent block: none,</li><li>. present and unmodified block: none,</li><li>. block present and modified: block purge request (whether this block is present or in the process of being transferred from the central memory RAM to the cache memory MC, state indicated by the bit -a- of the descriptor),</li></ul></li><li>shared data block write request:<ul id="ul0026" list-style="none"><li>. absent block: none,</li><li>. unmodified present block: invalidate the block,</li><li>. block present modified: block purge request (same remark as above),</li></ul></li><li>block informative write request:<ul id="ul0027" list-style="none"><li>. absent block: none,</li><li>. unmodified present block: invalidate the block,</li><li>. present block modified: impossible case,</li><li>. if an informative write request is pending on this same block, it is transformed into a block write request, because this block has been invalidated; for this purpose, the current request is canceled and an acknowledgment is returned to the request management processor PGU. The latter consults the RG directory and will find the missing block: a request is sent to the PGS management processor. This operation is taken into account by the parallel management processor PGP.</li></ul></li></ul>
If the signal <o ostyle="single">valid</o> is activated by the PGP parallel link management processor associated with the spy processor, then the latter's task is as follows, depending on the type of request (called local request):<ul id="ul0028" list-style="dash"><li>request to read a block of unshared data or request to write a block update: none,</li><li>shared data block read request: mark the valid block, awaiting transfer and not modified (v = a = 1 and m = 0),</li><li>shared data block write request: mark the valid block, awaiting transfer and modified (v = m = a = 1),</li><li>block informative write request: mark the block "modified" (m = 1).</li></ul>
To perform the functions described above, the spy processor PE includes a phase distributor DP_E, consisting of a shift register SR40 and a flip-flop B40. The output of a D40 differentiator circuit is connected to the input<o ostyle="single">S</o> from scale B40 and at the entrance <o ostyle="single">clear</o>40 SR40 register, as well as at the entrance <o ostyle="single">S</o> a B41 scale. The output 0 of the flip-flop B40 is connected to the input serial_in40 of the register SR40. The clk40 input of the SR40 register receives the general clock signal h and the validation input<o ostyle="single">en40</o> is connected to output 0 of flip-flop B43. Phase ϑ41, inverted, from register SR40 is connected to the input<o ostyle="single">R</o> of scale B40.
The operation of the phase distributor conforms to the description made in the PGU request management processor.
The signal <o ostyle="single">valid</o> of the BUSA bus is connected to the input of the differentiating circuit D40 and to the validation input <o ostyle="single">en41</o> a DEC41 decoder. The B41 scale receives on its input<o ostyle="single">R</o> the output of the logic gate NOU40. The exit<o ostyle="single">Q</o> of flip-flop B41, with open collector, is connected to the signal <o ostyle="single">done</o> from the BUSA bus.
The signals <o ostyle="single">valid</o> and <o ostyle="single">done</o> ensure the synchronization of the spy processor PE with the other spy processors of the multiprocessor system: the negative transition of the valid signal triggers the differentiator circuit D40 which produces a pulse allowing to activate the phase distributor and to put the signal <o ostyle="single">done</o> in the logic zero state via the flip-flop B41. The end of the spy's work is signaled by a change of state on the output of the logic gate NOU40, which produces the signal putting logic one<o ostyle="single">done</o> via the B41 scale.
The spy's job depends on the nature of the request and, for this purpose, the standard field of the BUSA bus is connected to the input of the DEC41 decoder. The dnp and dmaj outputs of the DEC41 decoder are connected to the inputs of an OU40 logic gate. The outputs dl, de, dei of the decoder DEC41 are connected to the inputs of a logic gate OU41, the outputs of and dei being also connected to the inputs of a logic gate OU42. The output of the logic gate OU40 is connected to one of the inputs of the logic gate NOU40 via a logic gate ET40, which receives on its other input the signal ϑ41. The output of logic gate OU41 is respectively connected to one of the inputs of logic gates ET42, ET43, which receive on their other input respectively the phase signal ϑ41 and ϑ44. The outputs of logic gates ET42 and ET43 are connected respectively to the inputs<o ostyle="single">S</o> and <o ostyle="single">R</o> a B44 scale. The output of logic gate ET42 is also connected to the inputs<o ostyle="single">S</o> of scales B42 and B43. A logic gate NOU41 receives on its inputs the phase signal ϑ45 and the output of a logic gate ET45, which receives on its two inputs the phase ϑ44 and the reverse output of a logic gate OU43. The output of the logic gate NOU41 is connected to the input<o ostyle="single">R</o> of scale B42. The exit<o ostyle="single">Q</o> of scale B42 is connected to the signal <o ostyle="single">rqst</o>_PR and the signal <o ostyle="single">grnt</o>_PR is issued upon entry <o ostyle="single">R</o> of flip-flop B43, to the inputs for validation of BV40 passage buffers, and to one of the inputs of logic gate OU44. The logic gate OU44 receives on its other input the output<o ostyle="single">Q</o> of a scale B45, and its output validates the buffers of passage BV41. The output of logic gate ET45 is also connected to one of the inputs of logic gate NOU40 which receives on another input the phase signal ϑ45. The output of the logic gate OU42 is connected to the inputs of logic doors ET46 and ET47, which also receive on their inputs the output of a logic gate OU43 and the phases ϑ44 and ϑ45 respectively. The exit<o ostyle="single">Q</o> of flip-flop B44 delivers the signal <o ostyle="single">r</o>_RG via a BV40 buffer, the output <o ostyle="single">Q</o> of flip-flop B45 delivers the signal <o ostyle="single">w</o>_RG via a BV40 buffer. The frame field of the common BUSA bus is connected to the adr_RG bus via a BV40 buffer. The data_RG bus is connected to the outputs of the BV41 validation buffers and to the input of the REG40 register, which receives on its load40 input the phase signal ϑ43. The REG40 register output is connected, for the tag part, to the inputs of the BV41 buffers and to one of the inputs of a COMP40 comparator. COMP40 comparator receives on its other input the tag part of the BUSA bus. The validation bit v, coming from the register REG40, connected to the comparator COMP40 has the constant value opposite 1. The bits v, a, m, f at the output of the register REG40 are respectively connected to one of the inputs of the multiplexers MUX40, MUX41, MUX42, MUX43. The outputs of these multiplexers provide the state of these same bits at the input of the BV41 buffers. The multiplexer MUX40 receives on its other inputs the constants zero and one, the inputs sel40 are connected to the output of an ET48 logic gate and to a signal <o ostyle="single">dlle</o>. The multiplexer MUX41 receives on its other input the constant one, selected when its input sel41, receiving a dlle signal from the management processor PGP, is in logic state one. The multiplexer MUX42 receives on its other inputs the constant one and the signal<o ostyle="single">r</o>_adrbloc_SR, its sel42 inputs receive signals <o ostyle="single">dlei</o> and <o ostyle="single">dlle</o> from the PGP management processor. The multiplexer MUX43 receives on its other input the constant one selected when its input sel43, connected to the output of an ET49 logic gate, is in logic state one. The logic gate ET49 receives on its inputs the output eg40 of the comparator COMP40, the signal f inverted and the signal m. The logic gate ET48 receives on its inputs the output eg40, the signal m inverted, the signal dlei and the output of the logic gate OU42. The logic gate OU43 receives on one of its inputs the signal eg40, and on its other input the signal dlle. The output of the logic gate ET49 is also connected to the load41 input of the FIFO queue, already described in the processor for managing the PGS serial link. The BUSA frame and tag fields are linked to the input of the FIFO queue. The signals<o ostyle="single">dlei</o>, <o ostyle="single">dlle</o> and <o ostyle="single">r</o>_adrbloc_SP come from the PGP parallel link management processor, which also receives the dei signal from the DEC41 decoder.
The operation of the assembly is as follows: activation of the signal <o ostyle="single">valid</o> initializes the phase distributor DP_E and validates the decoder DEC41 which produces the activation of an output depending on the standard information coding the nature of the request in progress on the common bus BUSA. The active output can be:<ul id="ul0029" list-style="dash"><li>dnp: request to read unshared data. The signal<o ostyle="single">done</o> is deactivated when phase ϑ41 occurs;</li><li>dmaj: block update write request. The signal<o ostyle="single">done</o> is deactivated in phase ϑ41;</li><li>d1: block read request;</li><li>of: block write request;</li><li>dei: informative writing request.</li></ul>
These three cases require read access to the RG directory, and the last two require possible rewriting of the directory. To this end, an access request is sent to the PGR directory management processor by flip-flop B42 (signal<o ostyle="single">rqst_PR</o>), flip-flop B43 inhibiting the phase distributor until access agreement, signified by the signal <o ostyle="single">grnt</o>_PR. Reading is therefore carried out from ϑ41 to ϑ44 (flip-flop B44) and possible writing from ϑ44 to ϑ45 (flip-flop B45) with storage of the descriptor in the REG40 register during phase ϑ43. If the block is absent from the MC cache memory (eq40 = 0), then the signal<o ostyle="single">done</o> is deactivated on phase ϑ44. If the block is present in the cache (eq40 = 1), then:<ul id="ul0030" list-style="dash"><li>if m = 1, a purge request is activated (activation of logic gate ET49) provided that this block is not already in the queue (f = 0); the only modified bit is f, set to one, when the descriptor is rewritten,</li><li>if m = 0, the block is disabled by the MUX40 multiplexer (activation of the ET48 logic gate),</li><li>if the demand is local (<o ostyle="single">dlle</o> or <o ostyle="single">dlei</o> active), then:<ul id="ul0031" list-style="none"><li>1) in the case of reading or writing, the bits v and a are set to 1 and m is set to 0 (read) or 1 (write) (signal state <o ostyle="single">r</o>_adrbloc_SP),</li><li>2) in the case of informative writing, the bit m is forced to 1.</li></ul></li></ul>
In these latter cases which require a rewrite in the RG management directory, the signal <o ostyle="single">done</o> is deactivated on phase ϑ45.
The processor for managing the parallel link PGP, an example of which is shown in FIG. 16, is responsible for requesting the common bus BUSA and carrying out the transaction requested, either by the request management processor PGU or by the management processor of the PGS serial link.
A request from the PGU request management processor can only be an informative write request. A request from the PGS serial link management processor is either a block read or write request or a block update request.
The processor for managing the parallel link PGP comprises a flip-flop B60, connected by its data input D to the signal <o ostyle="single">rqst</o>_UP. A flip-flop B61 is connected by its data input D to the signal<o ostyle="single">rqst</o>_SP. The Q and<o ostyle="single">Q</o> of a flip-flop B62 are connected respectively to the clock inputs of flip-flops B60 and B61. The exit<o ostyle="single">Q</o> of flip-flop B62 is looped back to its data input. The exits<o ostyle="single">Q</o> flip-flops B60 and B61 are connected to the inputs of an OU60 logic gate. The output of the logic gate OU60 is connected, on the one hand, to a differentiating circuit D60, on the other hand and inversely, to an input of a logic gate ET60, which receives on its other input the signal of general clock h. The output of the logic gate ET60 is connected to the clock input of the flip-flop B62. The output of the differentiator circuit D60 is connected to the input<o ostyle="single">S</o> a B63 scale. The exit<o ostyle="single">Q</o> of scale B63 outputs a signal <o ostyle="single">rqsti</o> to the referee AB, and his entry <o ostyle="single">R</o> is connected to the output of an ET62 logic gate. The signal<o ostyle="single">grnti</o> from the arbiter AB is connected to logic gates OU62 and NOU60. Logic gate OU62 receives the signal on its other input<o ostyle="single">valid</o> inverted, and its output is connected to the input of a differentiator circuit D61. The output of this circuit D61 is connected to one of the inputs of logic gate ET62 and to the input<o ostyle="single">S</o> of a B64 scale. The output Q of this flip-flop B64 is connected to the validation input of passing buffers BV60 and inversely, through an open collector inverter I60, to the signal<o ostyle="single">valid</o>. The signal<o ostyle="single">done</o> is connected to the input of a D62 differentiator circuit. The output of this circuit D62 is connected to the input<o ostyle="single">R</o> of flip-flop B64 and to one of the inputs of the NOU60 logic gate. The output of this NOU60 logic gate is connected to one of the inputs of NET60 and NET61 logic gates, which receive the outputs respectively on their other input.<o ostyle="single">Q</o> flip-flops B60 and B61. The exits<o ostyle="single">Q</o> flip-flops B60 and B61 are also connected respectively to the validation inputs of the passing buffers BV61 and BV62. The output of the NET60 logic gate is connected, on the one hand, to the input<o ostyle="single">S</o> from the flip-flop B60, on the other hand, to one of the inputs of a logic gate ET63, which receives on its other input the output of a differentiator circuit D63. ET63 logic gate output delivers signal<o ostyle="single">acq</o>_UP to the PGU management processor. The output of the logic gate NET61 is connected to the input<o ostyle="single">S</o> of flip-flop B61 and provides the signal <o ostyle="single">ack</o>_SP. The adr_bloc_UP bus is connected to the input of the BV61 validation buffers and to one of the inputs of a COMP60 comparator.
The adr_bloc_SP bus is connected to the input of the BV62 validation buffers. The outputs of the BV61 and BV62 buffers are connected together and to the input of the BV60 validation buffers. The output of the BV60 buffers is connected to the common BUSA bus. Logic gates OU63 and OU64 receive the outputs on their respective inputs<o ostyle="single">Q</o> flip-flops B60 and B61, the logic signal <o ostyle="single">grnti</o> and the signal mej for OU64. The output of the logic gate OU63 delivers the signal<o ostyle="single">dlei</o>, and the output of the logic gate OU64 the signal <o ostyle="single">dlle</o>. The other input of comparator COMP60 receives the tag and frame fields of the common bus BUSA. The input en60 of the COMP60 comparator is connected to the output of the logic gate ET61, which receives the signal on its inputs<o ostyle="single">grnti</o> inverted, the dei signal and the signal <o ostyle="single">rqst</o>_UP reversed. The output eg60 of comparator COMP60 is connected to the input of the differentiator circuit D63. The output of this circuit is also connected to the inputs<o ostyle="single">R</o> flip-flops B60 and B61 and at the other input of logic gate ET62.
The operation of the assembly is as follows: The flip-flops B60, B61 and B62 associated with the logic gates ET60 and OU60 constitute a local arbiter. This arbitrator examines the requests alternately<o ostyle="single">rqst</o>_UP and <o ostyle="single">rqst</o>_SP, and passes them on to the referee AB of the common bus BUSA by the signal <o ostyle="single">rqsti</o>. Access agreement is provided by signal validation<o ostyle="single">grnti</o> and the bus cycle takes place as soon as the signal is deactivated <o ostyle="single">valid</o> which releases the referee AB. Signal activation<o ostyle="single">done</o> releases the local arbitrator: the transaction is complete.
If the request comes from the PGS serial link management processor, then the signals <o ostyle="single">dlei</o> and <o ostyle="single">dlle</o> indicate to the associated spy processor PE the nature of the update of the state bits of the block to be carried out in the directory RG.
If the request comes from the request management processor PGU, then in the event of detection of informative writing on the same block (signal dei coming from the spy processor PE), an immediate release takes place: the request management processor PGU, after consulting the directory (the block has been invalidated) will request its request to the PGS serial link management processor.
The memory management processor PGM, an example of which is shown in FIG. 17, is responsible for ensuring the reading or writing of blocks in central memory RAM and participating in maintaining the consistency of the information in the various memories. MC covers of the multiprocessor system.
For this purpose, it includes a D80 differentiator circuit receiving on its input a signal <o ostyle="single">valid</o> and connected by its output to the inputs <o ostyle="single">S</o> of scales B80 and B81 as well as at the entrance <o ostyle="single">clr_80</o> an SR80 shift register. On the exit<o ostyle="single">Q</o> from flip-flop B80, with open collector, the signal is delivered <o ostyle="single">done</o>. The output Q of the flip-flop B81 is connected to the input serial_in80 of the register SR80; this flip-flop B81 is connected by its input<o ostyle="single">R</o> at output ϑ81, inverted, from register SR80. The register SR80 receives on its input clk80 the general clock signal h and its validation input<o ostyle="single">en80</o> is still active. The flip-flop B81 and the shift register SR80 constitute a phase distributor DP_M. The signal<o ostyle="single">valid</o> is also issued on the validation input <o ostyle="single">en81</o> a DEC80 decoder. This decoder is connected by its data input to the type part of the common bus BUSA, and provides the signals dnp, dl, de, ei and maj. A RAMFG 2-bit wide memory (called ro and rw respectively) receives on its address bus the tag and frame fields of the common bus BUSA. The data bus of this memory, consisting of the bits ro and rw, is connected on the one hand to a PAL80 logic, on the other hand to an ET80 logic gate, direct for rw, in reverse manner for ro. The PAL80 logic is connected to the standard field and receives logic signals cl, r / w, s / n, mff and en82: the signal cl comes from a B82 hascule, the signals r / w and s / n from a AFIFO associative queue, the mff signal from an ET81 logic gate, and the en82 signal from a B83 flip-flop, which receives on its inputs <o ostyle="single">S</o> and <o ostyle="single">R</o> respectively signals ϑ82 and ϑ81 inverted from the DP_M phase distributor. PAL logic cables on its ro-rw outputs the following logic equations: dnp = 10: dl.<o ostyle="single">mff</o> = 10; dl.mff = 01; de = 01; maj.<o ostyle="single">cl</o> = 00; ei = 01; key.cl.s /<o ostyle="single">not</o>.r / w = 10; key.cl.s /<o ostyle="single">not</o>.<o ostyle="single">r / w</o> = 01. The output of logic gate ET80 is connected to input D of a flip-flop B84, which receives phase horloge82 on its clock input. The output Q of this flip-flop is connected to one of the inputs of the logic gate ET81, which receives on its other input the output of the logic gate OU80. The two inputs of this OU80 logic gate are connected to the outputs of and dl of the DEC80 decoder. The read input r of the RAMFG memory is connected to phase ϑ81, and the write input w to the output of a logic gate ET82. The logic gate ET82 receives on its inputs the signal ϑ83 and the output of a logic gate ET83, whose inputs are connected to the signal s /<o ostyle="single">not</o> and in phase ϑ87. The dnp output of the DEC80 decoder is connected to logic gates ET84 and ET85, which receive on their other input respectively the phases ϑ81 and ϑ85. The signal s /<o ostyle="single">not</o> is also delivered to logic gates ET86 and ET87 which receive on their other input respectively the phases ϑ86 and ϑ90. The output mff of the logic gate ET81 is also connected to a logic gate ET88, which receives on its other input the phase ϑ83, and inversely to logic gates ET89 and ET90 which receive on their other input respectively the phases ϑ83 and ϑ87 . The output of logic gate ET88 is connected to the wff input of the AFIFO queue. The outputs of the logic gates ET84, ET86, ET89 are connected to the inputs of a logic gate OU81, the output of which is connected to the input S of a flip-flop B85. The outputs of the logic gates ET85, ET87, ET90 are connected to the inputs of a logic gate OU82, the output of which is connected to the input R of the flip-flop B85. The signal s /<o ostyle="single">not</o> is also connected, inversely, to a logic gate ET91 which receives on its other input the phase ϑ89. The output of logic gate ET91 is connected to the input of a logic gate NOU80 which receives on its other input the output of logic gate OU82. The output of the logic gate NOU80 is connected to the input<o ostyle="single">R</o> of scale B80. The maj output of the DEC80 decoder is connected to the input of logic gates ET92, ET93, ET94, ET95, which receive on their other input respectively the phases ϑ81, ϑ85, ϑ85, ϑ91. The outputs of the logic gates ET92 and ET93 reversed are connected respectively to the inputs S and R of a flip-flop B86, and those of the logic gates ET94 and ET95 to the inputs S and R of a flip-flop B82. The output Q of the flip-flop B82 produces the logic signal cl also supplied to the input cff of the AFIFO queue and to the command input sel80 of a multiplexer MUX80. The tag part, frame of the common bus BUSA is connected to the data input of the AFIFO queue and to one of the data inputs of the MUX80 multiplexer. The output d1 of the DEC80 decoder is also connected to one of the data inputs of the AFIFO file in order to produce the read / write signal 1 / e. The data output of the AFIFO file is connected to the other input of the MUX80 multiplexer. The output of the multiplexer MUX80 is connected to the address bus of the central memory RAM for the tag.cadre part, and the inputs of decoders DEC81 and DEC82 for the cpu field part. The exit<o ostyle="single">Q</o> of the flip-flop B86 is connected to the write input of the central memory RAM and to the input <o ostyle="single">en84</o> DEC82 decoder. The output Q of the flip-flop B85, slightly delayed, is connected to the reading input of the central memory RAM and to one of the inputs of a logic gate ET96, which receives on its other input the output of the logic gate OU82. The output of logic gate ET96 is connected to the input en83 of the decoder DEC81. The output j of the DEC81 decoder is connected to the validation input of the buffers of passage of the RDM memory shift register<sub>j</sub> and the output j of the decoder DEC82 at the loading input of said memory shift register RDM<sub>j</sub>.
The operation of this set is as follows: Signal activation <o ostyle="single">valid</o> causes the DP_M phase distributor to be triggered, and the DEC80 decoder to be validated, which will make it possible to determine the nature of the request. The phase ϑ81 is used to read the state of the bits corresponding to the block requested in the RAMFG memory, and the combination ro.rw is memorized in the flip-flop B84. A first write takes place in the RAMFG memory on phase ϑ83, which makes it possible to update the status bits. Their value is provided by the PAL80 logic and makes it possible to obtain the following sequences:<ul id="ul0032" list-style="dash"><li>in the case of a request for a block of unshared data (dnp) then whatever the state of the bits ro.rw (rw is necessarily zero), state 10 is forced ("block broadcast in read");</li><li>in the event of a block request in read (dl) or in write (de), if ro.rw = 01, then the request is queued on phase ϑ83 and state 01 is forced (in fact, this is the previous state), otherwise state 10 is forced in the event of reading ("block broadcast in reading") and state 01 is forced in the event of writing ("block broadcast in writing");</li><li>in the event of an update request (maj), state 00 is forced ("block not broadcast"). In these various cases, a reading or a writing in central memory RAM is operated, towards or from the memory shift register RDM<sub>j</sub> identified by the cpu field of the common bus BUSA. In the example chosen, the duration of the RAM memory cycle is 4 periods of the general clock h. In the case of reading of unshared data, the cycle is carried out from ϑ81 to ϑ85, in the other cases from ϑ83 to ϑ87. The writing is carried out from ϑ81 to ϑ85;</li><li>in the case of informative writing, this does not cause any movement of data, but forces the status bits to the value 01 (the starting state is in this case necessarily 10);</li><li>in the event of an update request, consultation of the AFIFO queue is systematically carried out. This consultation can lead to the reading of a block, in the case where a central processing unit CPU is waiting in the AFIFO queue to update this block.</li></ul>
The reading is carried out from ϑ86 to ϑ90 and the state of the bits is forced to 10 (request for a reading) or 01 (request for a writing). The end of any operation results in the reset of the flip-flop B80 which activates the signal<o ostyle="single">done</o>. This deactivation can occur on phases ϑ85, ϑ87 or ϑ91 depending on the requested operation, or on ϑ89 if the consultation of the queue gives a negative result.
The associative queue is not detailed. It is conventionally made up of an associative memory used in the queue. The number of words in this memory is equal to the number of central processing units of the multiprocessor system. An internal "daisy-chain" identifies on each phase ϑ81 the next candidate word for a write, which occurs if necessary on phase ϑ83 by the signal wff. The signal cff triggers a comparison from phase ϑ85, the flip-flops of the response memory having been reset to zero on phase ϑ84. The result of the comparison is reflected on the signal s /<o ostyle="single">not</o> (some /<o ostyle="single">none</o>) and the content of the concerned word is available on the data output from phase ϑ86. This word is then invalidated on phase ϑ90.
In the architecture described above, the spy processors PE<sub>j</sub> are requested at each address transfer on the common BUSA bus, with possible consultation of their MC cache memory<sub>j</sub>. This consultation is most of the time useless (low probability of presence of the address of the block corresponding to the transfer, in the cache memories).
It should be noted that the memory management processor PGM maintains status bits of the blocks and makes it possible to centralize management of maintaining consistency. To this end, it is possible to add to the architecture described above (FIG. 10) a parallel synchronization bus operating according to the same algorithm as the SYNCHRO synchronization bus of the variant which is described below. The spy processors are no longer strictly speaking spies (since they are connected to the synchronization bus and not to the common BUSA bus), and are designated by consistency maintenance processors (PMC<sub>j</sub> for the variant of figure 18). Thus, the memory management processor PGM remains requested at each transfer on the common bus BUSA, but the consistency maintaining processors are requested by the processor PGM only when they are affected by the transfer.
FIG. 18 presents a variant in which the consistency is maintained according to the principle mentioned above. This variant takes up the general architecture of FIG. 6, with block addresses which pass through the serial links LS<sub>j</sub>. This system includes a parallel SYNCHRO synchronization bus, with the same logical structure as the common bus BUSA, but controlled on the sole initiative of the memory management processor PGM.
The structure of the CPU<sub>j</sub> conforms to that shown in Figure 10, with some modifications:<ul id="ul0033" list-style="dash"><li>the structure of the MC cache memory<sub>j</sub> remains the same, as well as the RG management directory structure<sub>j</sub>,</li><li>the PGP parallel management processor<sub>j</sub> disappears, since the common bus BUSA no longer exists and the functions which were assigned to it are transferred to the management processor of the PGS serial link<sub>j</sub> ;</li><li>the spy processor PE<sub>j</sub> is replaced by a PMC consistency maintaining processor<sub>j</sub> which takes care of maintaining the status bits of the blocks in the cache-memory MC<sub>j</sub> to ensure consistency and which is activated on the sole initiative of the memory management processor PGM<sub>j</sub> via the SYNCHRO synchronization bus;</li><li>the PGU request management processor only knows one partner: the PGS serial link management processor<sub>j</sub>, to which he defers all his requests;</li><li>the PGS serial link management processor<sub>j</sub> is responsible for the transfer of addresses and data, in accordance with the principle described for the system of FIG. 6, each address being prefixed by the nature of the request;</li><li>the functionalities of the memory management processor PGM are those described with reference to FIG. 17, its activation being no longer ensured by the signal <o ostyle="single">valid</o>, which disappears (since previously associated with the common bus BUSA), but by the arbiter ABM described in the system of FIG. 6, which serializes the service requests which pass through the serial links. The RAMFG memory also consists of an additional field cpu associated with the status bits ro.rw.</li></ul>
The general operation of the embodiment shown in FIG. 18 is as follows: Each request from the processing processor CPU activates the request management processor PGU<sub>j</sub> with the indication read or write and code or data. This processor requires access to the RG management directory<sub>j</sub> with the PGR directory management processor<sub>j.</sub> Consulting the directory leads to one of the following cases:<ul id="ul0034" list-style="dash"><li>the block is present in the MC cache memory<sub>j</sub>, with the unmodified state (m = 0); if the request is a read, the requested information is extracted from the cache memory MC<sub>j</sub> and supplied to the processing processor CPU<sub>j</sub>. If the request is a write, then an informative write request ei is transmitted to the management processor of the PGS serial link.<sub>j</sub> ;</li><li>the block is present in the MC cache memory<sub>j</sub>, with the modified state (m = 1); the request, read or write, is satisfied;</li><li>the block is absent from the MC cache memory<sub>j</sub> ; a read or write block request is transmitted to the PGS serial link management processor<sub>j</sub>.</li></ul>
Thus, the requests made to the serial link management processor can be: a request to read unshared data (code): dnp, a request to read a block: d1, a request to read a block with a view to there write: from, an informative write request: ei.
To these various states, it is necessary to add the updated update state corresponding to the purging of a block, or at the request of the processor for maintaining consistency PMC<sub>j</sub>, or to free a block location in the cache memory. The addresses thus prefixed pass through the serial links LS<sub>j</sub>, and in accordance with the principle stated during the description of the architecture of FIG. 6, request the ABM arbiter when:<ul id="ul0035" list-style="dash"><li>in the case of a block read, the address is transmitted,</li><li>in the case of a block write, the address and the data are transmitted.</li></ul>
These requests are processed sequentially by the memory management processor PGM, with the same general structure as that described in FIG. 17. Their processing is as follows:<ul id="ul0036" list-style="none"><li>1 / dnp: request for unshared data. The block is transmitted and takes the state ro.rw = 10.</li><li>2 / dl: request to read a block. If the block is in the "not broadcast" (ro.rw = 00) or "broadcast in read" state (ro.rw = 10), it is transmitted and takes or keeps the ro.rw = 01 state. If the block is in the "broadcast in write" state, the request is put in the AFIFO queue. The memory management processor PGM then finds in the cpu field of the RAMFG memory the address of the cache memory MC<sub>i</sub> which contains the updated version of the requested block. A purge request is then sent on the SYNCHRO synchronization bus, only intended for the PMC consistency maintenance processor<sub>i</sub> associated with MC cache memory<sub>i</sub> concerned. This request can be qualified as an order addressed. Note that the PMC coherence maintaining processor<sub>i</sub> does not have to consult the RG management directory<sub>i</sub> associated since the memory management processor PGM is aware of the fact that it is the sole owner of the updated copy. Its role is simply to collect the request and deposit it in the FIFO queue<sub>i</sub> associated.</li><li>3 / de: request to read a block in order to write to it. If the block is in the "not broadcast" state (ro.rw = 00), it is transmitted and assumes the "broadcast in write" state (ro.rw = 01). If the block is in the "broadcast in read" state (ro.rw = 10), then the memory management processor issues a universal block invalidation command, then transmits the block with the "broadcast in write" state. (ro.rw = 01). Universal control activates all PMC coherence maintaining processors<sub>j</sub> which strictly perform the same operations as those described for the system in Figure 10. If the block is in the "broadcast in write" state, the request is queued AFIFO. As before, the memory management processor PGM issues a command addressed to the sole owner of the updated copy.</li><li>4 / shift: request to write a block following a purge. The operating algorithm in this case is strictly the same as that described with reference to FIG. 17 for the PGM processor. It should be noted that the problem of acknowledgment of writing naturally finds its solution in this embodiment by a command addressed to acknowledgment.</li><li>5 / ei: informative writing. This case is treated directly on the common bus BUSA in the architecture presented in FIG. 10. In the embodiment referred to here, and in order to guarantee synchronization, this operation is supported by the memory management processor PGM. If the block is in the "broadcast in read" state, then a command, both universal and addressed, is issued: addressed in the sense that the PMC processor<sub>j</sub> concerned notes the acknowledgment of the informative writing request and passes the block concerned in the "modified" state in the RG management directory<sub>j</sub> , universal in the sense that all other PMC processors<sub>i</sub> must invalidate this block in their directory.</li></ul>
The block in the "broadcast in write" state indicates that an informative write request has been processed on this same block during the waiting time for processing the request. In this case, the informative writing request is transformed into a writing request from, and follows the same processing as in the corresponding writing case.
The SYNCHRO synchronization parallel bus is responsible for broadcasting block addresses prefixed by a processor number and a request type, ie approximately 30 to 40 bits depending on the characteristics of the multiprocessor. This information is also transmitted unidirectionally. Their transfer can then advantageously again be done by a serial link. The transfer rate is less critical than for the blocks, and simplified solutions can be envisaged, for example by means of "TAXI" circuits manufactured by the company "AMD".
FIG. 19 presents a partial block diagram of an architecture according to the invention, in which several central processing units UC<sub>k</sub>... are united in a cluster and share the same LS serial link<sub>k</sub>. For this purpose, a local ABL referee<sub>k</sub> associated with the cluster is responsible for arbitrating access conflicts to block address communication means, and sharing with the PGM memory processor, fitted for this purpose, a busy signal<sub>k</sub> permanently indicating the free or occupied state of the LS serial link<sub>k</sub>. Means for coding and decoding an identification header of the processor concerned within a cluster are associated with the logic for transmitting and receiving the data blocks.
In the case where the block address communication means are constituted by the common bus BUSA, the operation is as follows: If the central processing unit CPU<sub>k + j</sub> wish to make a block transfer in the direction of central memory RAM to cache memory MC<sub>k + j</sub> (case dnp, dl, de) or carry out an informative writing ei, then it requires access to the common bus BUSA to the local arbiter ABL<sub>k</sub>, who passes the request on to the referee AB. The BUSA common bus access agreement is returned to the central processing unit CPU<sub>k + j</sub> and the transfer is carried out as described with reference to FIG. 10. Any block transmitted in the direction from central memory RAM to cache memory MC<sub>k + j</sub> must then be identified, because the order of the requests is not respected due to the possibility of queuing in the AFIFO queue of the memory management processor PGM. If the central processing unit CPU<sub>k + j</sub> want to do a block transfer in the memory cache MC direction<sub>k + j</sub> to the central memory RAM (major case), then it first requires the local arbiter ABL<sub>k</sub> access to the LS serial link<sub>k</sub>. ABL local referee<sub>k</sub> and the memory management processor PGM are both capable of taking over the serial link LS<sub>k</sub> : the contention is avoided by synchronizing the modification of the busy signal<sub>k</sub> with the valid signal (the PGM memory management processor can only initiate or restart a transfer during a memory transaction). LS serial link occupancy agreement<sub>k</sub> drives CPU<sub>k + j</sub> to transfer its information block to the RDM memory shift register<sub>k</sub>, then to request from the local ABL arbitrator<sub>k</sub> access to the common bus BUSA in order to carry out the update request there, which is carried out according to the algorithm described with reference to FIG. 10. The writing of update can result in a block release in the AFIFO queue of the memory management processor PGM and request an RDM shift register<sub>j</sub> busy. In this case, the requested transfer is delayed and chained to the current transfer.
In the case where the means of communication of block addresses are the serial links themselves, the operation is then identical to the previous case with regard to the preemption of the serial link, and identical for the general operating algorithm to that presented with reference to FIG. 17.
For example, a block read request from the central processing unit CPU<sub>k + j</sub> first requires an access agreement to the LS serial link<sub>k</sub>, agreement given by local referee ABL<sub>k</sub> in consultation with the memory management processor PGM. The access agreement results in the transfer of the address of the requested block to the LS serial link<sub>k</sub>, which is immediately released: it is available for any other transaction if necessary. A block write request follows the same protocol for access to the LS serial link<sub>k</sub>.
In the architectures described with reference to FIGS. 1 to 19, there were as many RDM memory shift registers<sub>j</sub> than CPU central units<sub>j</sub> : a LS serial link<sub>j</sub> was statically assigned to a couple (RDM<sub>j</sub>, CPU<sub>j</sub>).
If there must obviously be at least one LS serial link<sub>j</sub> between a central unit and the central memory RAM, the number of shift registers RDM<sub>j</sub> may be less. Indeed, if tacc is the access time to the central memory RAM and ttfr the transfer time of a block, it is not possible to keep more than n = ttfr / tacc shift registers simultaneously occupied. For example, for tacc = 100 ns and ttfr = 1,200 ns, we get n = 12.
Tacc and ttfr are then criteria characteristic of the performance of the multiprocessor system according to the invention and the establishment of n registers with memory shift RDM<sub>j</sub> is compatible with a higher number of LS serial links<sub>j</sub> that on the condition of interleaving an RI interconnection network type logic between registers and links, the assignment of an RDM memory register<sub>j</sub> to a serial link LS<sub>j</sub> being performed dynamically by the memory management processor PGM.
Furthermore, the central memory RAM will generally consist of m memory banks RAM₁, ... RAM<sub>p</sub>, RAM<sub>m</sub> arranged in parallel, each memory bank comprising n RDM shift registers<sub>j</sub> connected by an RI interconnection network<sub>p</sub> to all LS serial links<sub>j</sub>. It is then possible, provided that the block addresses are uniformly distributed on the RAM memory banks<sub>p</sub>, to obtain a theoretical performance of mxn simultaneously active shift registers. The uniform distribution of addresses is ensured by conventional mechanisms of interleaving of addresses.
In FIG. 20a, an architecture according to the invention is partially represented, comprising m RAM memory banks<sub>p</sub> with n RDM shift registers<sub>j</sub> by memory banks and q CPU central units<sub>j</sub>. Each memory bank is of the random access type with a data input / output of width corresponding to a block of bi information, this input / output being (as previously for the RAM memory) connected by a bus parallel to the set of RDM elementary registers<sub>1p</sub>... RDM<sub>jp</sub>.
The interconnection network is of known structure ("cross-bar", "delta", "banyan" ...). It will be noted that a multi-stage network is well suited insofar as the path establishment time is negligible compared to its occupation time (the transfer time of a block) and that it only concerns one bit per link.
The memory management processor PGM is adapted to be able to dynamically allocate an output to an input of the network, that is to say to connect a memory shift register RDM<sub>j</sub> and a serial link LS<sub>i</sub>.
In the case where the block address communication means are constituted by the common bus BUSA, the operation is as follows: In the event of a request to read a block from the central processing unit CPU<sub>j</sub>, the memory management processor PGM<sub>p</sub> concerned allocates an RDM shift register<sub>i</sub>, controls the RI interconnection network accordingly and initiates the transfer.
In the event of a request to write a block from the central processing unit CPU<sub>j</sub>, a path must first be established. To this end, a first path establishment request is sent on the common bus BUSA, followed by the effective write request upon transfer of the block from the cache memory MC<sub>j</sub> to the RDM shift register<sub>i</sub>. During the first request, the memory management processor PGM is responsible for allocating a path and controlling the interconnection network RI.
In the case where the block address communication means are the serial links themselves, a path must be established prior to any transfer. This problem is identical to the classic problem of sharing a set of n resources by m users and can be solved by conventional solutions of arbitration of access conflicts (communication protocols, additional signals).
In the example architecture shown in Figure 5, the RDM shift registers<sub>j</sub> and RDP<sub>j</sub>, their validation logic LV1 and LV2 were carried out using rapid technology, the assembly being synchronized by a clock of frequency F at least equal to 100 MHz.
FIG. 20b presents in alternative to the architecture proposed in FIG. 20a, a solution according to the invention in which each LS serial link<sub>j</sub>, which connects the CPU processor<sub>j</sub> to all memory banks, is split into m LS series links<sub>jp</sub> connecting point to point the CPU processor<sub>j</sub> to each of the RAM memory banks<sub>p</sub>.
This process has the following double advantage:<ul id="ul0037" list-style="dash"><li>each link being of the point to point type, can be better adapted from the electrical point of view or from the optical fiber point of view,</li><li>an additional level of parallelism is obtained when the processing processor is able to anticipate block requests, which is currently the case for the most efficient processors.</li></ul>
Interface logic (previously noted TFR<sub>j</sub> and RDP<sub>j</sub>) which is associated with the LS serial link<sub>j</sub>, processor side CPU<sub>j</sub>, is then duplicated in m copies I₁ ... I<sub>p</sub>... I<sub>m</sub>. Note the presence of a link to maintain information consistency, private to each RAM memory bank<sub>p</sub>. The operation of this link is similar to that of the SYNCHRO bus in Figure 18.
FIGS. 21a and 21b show another RAM memory structure which includes 2<sup>u</sup> memory plans, each memory plan having a front of t / 2<sup>u</sup> binary information (for reasons of clarity, we have shown in FIG. 21a the means necessary for reading a bi block, and in FIG. 21b the means necessary for writing). RDM shift registers<sub>j</sub> or RDP<sub>j</sub> consist of 2<sup>u</sup> RDM elementary shift sub-registers<sub>jp</sub> at t / 2<sup>u</sup> capacity bits. The example presented in FIG. 20 is an embodiment with 8 memory planes (u = 3). (For clarity of the drawing, a single RDM shift register has been shown<sub>j</sub> formed for all RDM sub-registers<sub>jp</sub>). Each RAM memory plan<sub>p</sub> has a set of RDM elementary shift registers facing its access front<sub>jp</sub> and is capable of operating with an offset frequency of at least F / 2<sup>u.</sup>
The operation of the assembly in the event of reading is illustrated in FIG. 21a. A block is read synchronously in all 2<sup>u</sup> memory plans and loaded the same in elementary registers of the same rank. The serial outputs of these registers are connected to the inputs of a MUXR multiplexer produced in rapid technology (ASGA). A circuit of this type, perfectly adapted, is available from "GIGABIT LOGIC", under the reference "10GO40", and is capable of delivering a logic signal at a frequency of 2.7 GHz. It also provides a frequency clock divided by eight, which constitutes the RDM elementary register shift clock.<sub>jp</sub>.
In the case of writing, symmetrical operation, presented in FIG. 21b, is obtained with a DMUXR multiplexer circuit from the same manufacturer (referenced "10G41"), with the same performance characteristics.
A transfer frequency of 500 MHz is thus obtained with 8 elementary registers operating at a frequency of 500/8 = 62.5 MHz, which makes them achievable in more conventional technology ("MOS" for example).
The multiplexer and demultiplexer circuits referenced above can be combined into sets of 16, 32, ... bits. Thus, by associating respectively 16, 32 memory planes operating at 62.5 MHz, it is possible to obtain bit rates of 1 and 2 GHz, ie a performance level 2 to 4 times higher.
It should be noted that the TFR logic can be performed on one of the elementary registers, and that the validation logic LV is integrated into the "ASGA" circuits (open collector output).
FIG. 22 shows the general structure of a component of the “VLSI” integrated circuit type, called “serial multiport memory” and capable of equipping a multiprocessor system according to the invention. This component can be used in the multi-processor architecture described above, either to create the central memory RAM and the shift registers RDM<sub>j</sub> associated, either to create each MC cache memory<sub>j</sub> and its RDP shift register<sub>j</sub>. To simplify the notations, the following description has kept the symbols relating to the central memory RAM and the associated shift registers.
The list of pins of this circuit with the corresponding signals is as follows:<ul id="ul0038" list-style="dash"><li>adbloc₀-adblocm<sub>m-1</sub> : m bits of bi block addresses,</li><li>admot₀-admot<sub>k-1</sub> : k bits of word addresses in the block,</li><li>numreg₀-numreg<sub>n-1</sub> : n bits of rd register addresses,</li><li><o ostyle="single">cs</o> : "chip select": circuit selection signal,</li><li><o ostyle="single">wr</o> : "write": write signal,</li><li><o ostyle="single">rd</o> : "read": read signal,</li><li>bit/<o ostyle="single">block</o> : multiport function control signal,</li><li>normal/<o ostyle="single">config</o> : operating mode signal,</li><li>data₀-data₁₋₁: 1 bits of data,</li><li>h₁-h<sub>not</sub> : n clock signals,</li><li>d₁-d<sub>not</sub> : n data signals.</li></ul>
The values m, n, l are functions of the current state of technology. Current values could be:<ul id="ul0039" list-style="dash"><li>m = 16 or 2¹⁶ bi blocks of 64 bits each (i.e. 4 Mbits),</li><li>n = 3, i.e. 8 rd registers,</li><li>1 = 8 or a parallel byte type interface,</li><li>k = 3 due to the presence of 8 bytes per block.</li></ul>
The target component has around fifty pins.
This serial multiport memory circuit is composed of a random access random access memory RAM, of predetermined width t, capable of being write-controlled on independent edges of width t / 4 (value chosen by way of example in FIG. 22 ) and t / 1. The data lines of this RAM memory are connected to the inputs of a “barrel shifter”, or MT multiplexing, depending on the component version, the MT multiplexing logic being able to be considered as offering a subset of the possibilities of the "barrel" logic and therefore simpler to carry out. The address and control signals of this RAM memory, namely csi, wri, rdi, adbloci, are delivered from a COM control logic. This COM logic also receives the information signals from the pins <o ostyle="single">cs</o>, <o ostyle="single">wr</o>, <o ostyle="single">rd</o>, bit /<o ostyle="single">block</o>, normal /<o ostyle="single">config</o>, numreq and is connected, on the one hand, by "format" command lines to the "barrel" type logic BS, on the other hand, to the output of a configuration register RC1, and to the input LSR selection logic providing srd₀, ... srd signals<sub>n-1</sub> and src₁, src₂, src₃. The outputs of the barrel type logic BS constitute an internal bus for parallel communication BUSI, connected to a set of shift registers RD₀, ... RD<sub>n-1</sub>, on the one hand, on their parallel inputs and, on the other hand, on their parallel outputs through validation buffers BV100₀, ... BV100<sub>n-1</sub> and at the parallel input of the configuration registers RC₁, RC₂ ... RC<sub>i</sub>.
The 1 least significant bits of the BUSI bus are also received on the 1 data₀ ... data₁₋₁ pins. Each RD shift register<sub>i</sub> and associated logic gates constitute an ELRD functional unit<sub>i</sub>, driven by a set of logic elements which constitute an LF forcing logic<sub>i</sub>. Each ELRD functional unit<sub>i</sub> has ET100 logic gates<sub>i</sub> and ET101<sub>i</sub>'' connected on one of their input to the srd output<sub>i</sub> of the LSR selection logic, and receiving on their other input respectively the rd signals<sub>i</sub> and wr<sub>i</sub>. ET100 logic gate output<sub>i</sub> is connected to the load100 input<sub>i</sub> of the shift register RD<sub>i</sub>, as well as at the load101 entry<sub>i</sub> and at the input S respectively of a CPT100 counter<sub>i</sub> and a B100 scale<sub>i</sub> belonging to the LF forcing logic<sub>i</sub>. ET101 logic gate output<sub>i</sub> is connected to the control input of the BV100 validation buffers<sub>i</sub>. The output di is connected to the output of a logic gate PL<sub>i</sub>, which receives the serial output of the shift register RD on its data input<sub>i</sub> and on its command input the output of a logic gate OU100<sub>i</sub>. The signal from pin h<sub>i</sub> is issued at entry clk100<sub>i</sub> of the RD register<sub>i</sub> as well as at the entry down100<sub>i</sub> of the CPT100 counter<sub>i</sub>. The zero100 output<sub>i</sub> of the CPT100 counter<sub>i</sub> is connected to the input R of the flip-flop B100<sub>i</sub>.
The LF forcing logic<sub>i</sub> additionally includes a MUX100 multiplexer<sub>i</sub> which receives the values t and t / 4 on its data inputs. The data output of the MUX100 multiplexer<sub>i</sub> is connected to the data input of the CPT100 counter<sub>i</sub>, and the sel100 selection command<sub>i</sub> of the MUX100 multiplexer<sub>i</sub> is connected to output 1 of the RC₁ register. The output Q of flip-flop B100<sub>i</sub> is connected to one of the inputs of an ET102 logic gate<sub>i</sub>, which receives on its other input the signal from pin i of a register RC₂. ET102 logic gate output<sub>i</sub> is connected to one of the inputs of the logic gate OU100<sub>i</sub>, which receives on its other input the signal from pin i of a register RC₃. The inputs for loading the registers RC₁, RC₂, RC₃ respectively receive the signals src₁, src₂, src₃ from the selection logic LSR.
This component has a dual function: if the bit /<o ostyle="single">block</o> is in the "bit" state, then the operation of this component is that of a conventional semiconductor memory: the adbloc signals associated with the admot signals constituting the address bus in word unit (8 bits in the example ), the signals <o ostyle="single">cs</o>, <o ostyle="single">rd</o>, <o ostyle="single">wr</o> have the usual meaning assigned to these signals, and the data pins carry the data.
Internally, when reading the information block designated by adbloc is read in RAM memory and presented to the input of the barrel BS or multiplexing logic MT. The combination of the admot and bit / signals<o ostyle="single">block</o> allow the COM control logic to supply the barrel BS or multiplexing MT logic with the "format" signals. The word concerned is then framed on the right at the output of the barrel or multiplexing logic and thus presented on the data pins.
Internally, in writing, the word presented on the data lines of data is framed by the barrel logic LS or by multiplexing MT by the same format control signals as in reading, with regard to its position in the block. The COM control logic then sends a partial wri write signal on the only "section" of memory concerned, and at the address designated by the adbloc signals.
If the bit / signal<o ostyle="single">block</o> is in the "block" state, then operation depends on the state of the normal signal /<o ostyle="single">config</o>. The config mode programs the configuration registers RC₁, RC₂, RC₃ addressed by the signals numreg, and programmed from the data lines data. The register RC₁ makes it possible to modify the size of the block: t and t / 4 in the example, that is to say 64 bits and 16 bits. Internally, the operation is similar to that described in the "bit" operating mode: t or t / 4 bits are framed on the internal bus BUSI (for reading), or opposite the "section" of the block concerned (in writing). Multiple block sizes can be considered (t, t / 2, t / 4 ...).
The register RC₃ makes it possible to choose for each register a permanent direction of operation: either at the input (RC3<sub>i</sub> = 0) or at output (RC3<sub>i</sub> = 1). This permanent direction makes it possible to adapt the component to serial links with permanent unidirectional links. The register RC₂ makes it possible to choose for each register, provided that the corresponding bit of RC₃ is in logic state 0, an operating mode with alternating bidirectional links: on a RAM memory read, the shift register RD<sub>i</sub> concerned "goes" in output mode for the time of transmission of the block, then returns to the idle state in "input" mode. Internally, the B100 scale<sub>i</sub>, which controls the logic gate PL<sub>i</sub>, is set to a load signal from the RDM register<sub>i</sub> and reset to zero at the end of the transfer of t or t / 4 bits, via the counter CPT100<sub>i</sub>, initialized at t or t / 4 depending on the state of the register RC₁, and which receives on its down counting input the clock pulses hi. In normal operation (normal signal /<o ostyle="single">config</o> in the normal state) for a reading, the block addressed by the adbloc pins is loaded in the register RD<sub>i</sub> addressed by pins numreg. If the block is partial (t / 4), then it is transmitted in the low weight position on the internal bus BUSI by the barrel type BS or MT multiplexing logic. This block is then transmitted upon activation of the clock signal hi.
In normal operation for a write, the content of the RD register<sub>i</sub> addressed by pins numreg is written in the adblock address RAM memory block. If the block is partial, it is transmitted in the most significant position on the internal BUSI bus, then framed opposite the section of the block concerned by the BS barrel or MT multiplexing logic, and finally a partial write signal. wri is issued on the affected section.
It should be noted that if a partial block is in service, then the address of this partial block in the block is provided by the address lines admot.
This component is perfectly suited to the various architectural variants described. Associated in parallel, 8, 16 ... circuits of this type make it possible to produce the device described in Figures 20a, 20b. If the RAM memory is in rapid technology, then this component can also be used at the cache memory level, by multiplexing, according to the device described in FIGS. 20a, 20b, the internal registers of the same component.
26 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26
Every citation, both waysCites: the store holds 4 of 5
| Document | Relation | Office |
|---|---|---|
| EP0126976A | Cites | European Patent Office (EPO) |
| EP0166192A | Cites | European Patent Office (EPO) |
| EP0187289A | Cites | European Patent Office (EPO) |
| WO8202615A | Cites | World Intellectual Property Organization (WIPO) |
| IEEE Transactions on Computers, vol. C-31, no. 11, novembre 1982, IEEE, (New York, US), M. Dubois et al.: "Effects of cache coherency in multiprocessors", pages 1083-1099. | Non-patent | – |
19 members in 6 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 8718103 | France | A | |
| 8718103 | France | – | |
| 8800608 | France | W | |
| 8718103 | – | – | – |
| FR19870018103 | – | – | – |
| FR8800608 | – | – | – |
| WO1988FR00608 | – | – | – |
Members19
| Document | Office | Kind | |
|---|---|---|---|
| FR2624631A1 | France | A1 | |
| WO8906013A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP0346420A1 | European Patent Office (EPO) | A1 | |
| FR2624631B1 | France | B1 | |
| JPH02502496A | Japan | A | |
| EP0346420B1This record | European Patent Office (EPO) | B1 | |
| DE3850770D1 | Germany | D1 | |
| DE3850770T2 | Germany | T2 | |
| US5598554A | United States of America | A | |
| US6112287A | United States of America | A | |
| US6345321B1 | United States of America | B1 | |
| US2002124153A1 | United States of America | A1 | |
| US2003018880A1 | United States of America | A1 | |
| US2003120895A1 | United States of America | A1 | |
| US6748509B2 | United States of America | B2 | |
| US2004133729A1 | United States of America | A1 | |
| US2004139285A1 | United States of America | A1 | |
| US2004139296A1 | United States of America | A1 | |
| US7136971B2 | United States of America | B2 |
28 legal events, as 3 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Notification of lapseLapsedST | ST | FR | |
| Nl: lapsed or anulled due to non-payment of the annual feeLapsedNLV4 | NLV4 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Gb: european patent ceased through non-payment of renewal feeCeasedGBPC | GBPC | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Transmission of propertyTP | TP | FR | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| European patent in force as of 2002-01-01IF02 | IF02 | GB | |
| No opposition filedOpposition26N | 26N | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Gb: translation of ep patent filed (gb section 77(6)(a)/1977)GBT | GBT | EP | |
| It: translation for a ep patent filedITF | ITF | EP | |
| It: translation for a ep patent filedITF | ITF | EP | |
| Corresponds to:REF | REF | EP | |
| Designated contracting statesAK | AK | EP | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 0346420
- Publication, DOCDB
- 0346420
- Publication, EPODOC
- EP0346420
- Application
- 89900274
- Application, DOCDB
- 89900274
- Application, EPODOC
- EP19890900274
Titles6
- German
- VERFAHREN ZUM INFORMATIONSAUSTAUSCH IN EINEM MEHRPROZESSORSYSTEM
- English
- PROCESS FOR EXCHANGING INFORMATION IN A MULTIPROCESSOR SYSTEM
- French
- PROCEDE D'ECHANGE D'INFORMATION DANS UN SYSTEME MULTIPROCESSEUR
- German
- VERFAHREN ZUM INFORMATIONSAUSTAUSCH IN EINEM MEHRPROZESSORSYSTEM.
- English
- PROCESS FOR EXCHANGING INFORMATION IN A MULTIPROCESSOR SYSTEM.
- French
- PROCEDE D'ECHANGE D'INFORMATION DANS UN SYSTEME MULTIPROCESSEUR.
Classification
- CPC, 1
- G06F12/0813
- IPC, 4
- G06F12 08
- G06F12 0813
- G06F15 16
- G06F15 177
Designated states1
- Contracting states, 1
- Netherlands (Kingdom of the)
