Memory buffers for merging local data from memory modules
Abstract
This record has no abstract on file.
Term
Term ended
Expired 27 January 2026, 0.7 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
6 claims: 6 independent, 0 dependent
- 1一以上のレーンを有するシリアル入出力インタフェースを備え、前記一以上のレーンのそれぞれは、 ローカル・データバスに結合したパラレル入力と、第1のクロック信号に結合されたクロック入力と、ロード信号に結合されたロード/シフトバー入力と、を有し、前記ローカル・データバス上のパラレルデータを第1のシリアル出力上のシリアル化されたローカルデータにシリアル化する第1のパラレル入力シリアル出力(PISO)シフトレジスタと、 前記第1のシリアル出力に結合された第1のデータ入力と、フィードスルー・データを受信するための第2のデータ入力と、ローカル・データ選択信号に結合された選択入力と、を有し、前記ローカル・データ選択信号に応じて前記シリアル化されたローカルデータとフィードスルー・データとをマージして多重化された出力上のシリアル・データストリームにする第1のマルチプレクサと、 前記シリアル・データストリームを受信するための前記多重化された出力に結合された入力を有し、シリアル・データリンク上へ前記シリアル・データストリームをドライブするトランスミッタと、 前記第1のマルチプレクサ及び前記第1のPISOシフトレジスタに結合され、前記第1のクロック信号及びマージ・イネーブル信号を受信し、前記マージ・イネーブル信号及び前記第1のクロック信号に応じて、前記シリアル化されたローカルデータ及び前記フィードスルー・データを前記シリアル・データストリームにマージする前記ローカル・データ選択信号を生成する制御ロジックと、 前記ローカル・データバス及び前記第1のPISOシフトレジスタの間で結合されたバス・マルチプレクサと、 を備え、 前記バス・マルチプレクサは、前記ローカル・データバスの複数のビットの半分に結合された第1のデータ入力と、前記ローカル・データバスの前記複数のビットの残り半分に結合された第2のデータ入力と、前記第1のPISOシフトレジスタのパラレル入力に結合された多重化出力と、データバス選択信号に結合された選択入力と、を有し、前記バス・マルチプレクサは、データバス選択信号に応じて、前記ローカル・データバスの前記複数のビットの前記半分又は前記ローカル・データバスの前記複数のビットの前記残りの半分に結合させて、前記第1のPISOシフトレジスタの前記パラレル入力にし、 前記制御ロジックは、前記マージ・イネーブル信号に結合された第1のデータ入力及びロード信号に結合された選択入力を有する、第2のマルチプレクサと、 前記第2のマルチプレクサに結合されたDタイプフリップフロップと、を含むマージ制御ロジックを有し、前記Dタイプフリップフロップは、前記第2のマルチプレクサの出力に結合されたデータ入力と、前記第1のクロック信号に結合されたクロック入力と、前記第1のマルチプレクサの前記選択入力及び前記第2のマルチプレクサの第2のデータ入力に結合されたデータ出力と、を有し、前記Dタイプフリップフロップは、前記データ出力上の前記ローカル・データ選択信号を生成するために、前記ロード信号及び前記第1のクロック信号に応じて、マージ・イネーブル信号を記録し、 前記第2のマルチプレクサは、前記ロード信号の反転信号に応じて、前記ローカル・データ選択信号を前記Dタイプフリップフロップの前記データ入力に再循環させる 集 積回路。
- 2一以上のレーンを有するシリアル入出力インタフェースを備え、前記一以上のレーンのそれぞれは、 ローカル・データバスに結合したパラレル入力と、第1のクロック信号に結合されたクロック入力と、ロード信号に結合されたロード/シフトバー入力と、を有し、前記ローカル・データバス上のパラレルデータを第1のシリアル出力上のシリアル化されたローカルデータにシリアル化する第1のパラレル入力シリアル出力(PISO)シフトレジスタと、 前記第1のシリアル出力に結合された第1のデータ入力と、フィードスルー・データを受信するための第2のデータ入力と、ローカル・データ選択信号に結合された選択入力と、を有し、前記ローカル・データ選択信号に応じて前記シリアル化されたローカルデータとフィードスルー・データとをマージして多重化された出力上のシリアル・データストリームにする第1のマルチプレクサと、 前記シリアル・データストリームを受信するための前記多重化された出力に結合された入力を有し、シリアル・データリンク上へ前記シリアル・データストリームをドライブするトランスミッタと、 前記ローカル・データバス及び前記第1のPISOシフトレジスタの間で結合されたバス・マルチプレクサと、 前記ローカル・データバスの複数のビットの残りの半分に結合されたパラレル入力と、前記第1のクロック信号に結合されたクロック入力と、前記ロード信号に結合されたロード/シフトバー入力と、前記第1のPISOシフトレジスタのシリアル入力に結合した第2のシリアル出力と、を含み、前記ローカル・データバスの前記複数のビットの前記残りの半分上のパラレルデータをシリアル化して、前記第1のPISOシフトレジスタの前記シリアル入力に結合するための前記第2のシリアル出力上のシリアル化されたローカルデータにする第2のパラレル入力シリアル出力(PISO)シフトレジスタ と、 を備え、 前記バス・マルチプレクサは、前記ローカル・データバスの複数のビットの半分に結合された第1のデータ入力と、前記ローカル・データバスの前記複数のビットの残り半分に結合された第2のデータ入力と、前記第1のPISOシフトレジスタのパラレル入力に結合された多重化出力と、データバス選択信号に結合された選択入力と、を有し、前記バス・マルチプレクサは、データバス選択信号に応じて、前記ローカル・データバスの前記複数のビットの前記半分又は前記ローカル・データバスの前記複数のビットの前記残りの半分に結合させて、前記第1のPISOシフトレジスタの前記パラレル入力にし、 前記第1のPISOシフトレジスタのパラレル入力は、前記ローカル・データバスの複数のビットの半分に結合される 集 積回路。
- 3前記フィードスルー・データは、2ビット長であり、 前記第1のPISOシフトレジスタのパラレル入力は、少なくとも6ビット長で、前記第1のPISOシフトレジスタのシリアル出力は、2ビット長であり、 前記第1のマルチプレクサは、2ビット・バス・マルチプレクサであり、前記多重化された出力の前記シリアル・データストリームは2ビット長であり、 前記トランスミッタは、前記2ビット長のシリアル・データストリームを受信して、それを前記シリアル・データリンク上へのシングル・ビットシリアル・データストリームとしてシリアル化する 請求項1 または2 に記載の集積回路。
- 4前記一以上のレーンのそれぞれは、 前記 第1の マルチプレクサ及び前記第1のPISOシフトレジスタに結合された制御ロジックを更に有し、前記制御ロジックは、前記第1のクロック信号と、モード選択信号と、マージ・イネーブル信号と、を受信し、前記制御ロジックは、前記マージ・イネーブル信号と前記第1のクロック信号に応じて、前記シリアル化されたローカルデータと前記フィードスルー・データとをマージさせて、前記シリアル・データストリームにするための前記ローカル・データ選択信 号を 生成し、前記制御ロジックは、前記モード選択信号に応じて、前記データバス選択信号を更に生成する、 請求項 2 に記載の集積回路。
- 5前記ロード信号は、前記第2のPISOシフトレジスタのロード/シフトバーのバー入力に結合された初期のロード・パルス信号であり、 前記制御ロジックは、マージ制御ロジックを有し、 前記マージ制御ロジック は、 前記マージ・イネーブル信号に結合された第1のデータ入力と、前記初期のロード・パルス信号に結合された選択入力と、を含む第2のマルチプレクサと、 前記第2のマルチプレクサに結合された第1のDタイプフリップフロップと、を含み、前記第1のDタイプフリップフロップは、前記第2のマルチプレクサの出力に結合されたデータ入力と、前記第1のクロック信号に結合されたクロック入力と、前記第1のマルチプレクサと前記第2のマルチプレクサの第2のデータ入力との前記選択入力に結合されたデータ出力と、を有し、前記第1のDタイプフリップフロップは、前記初期のロード・パルス信号と前記データ出力上の前記ローカル・データ選択信号を生成するための前記第1のクロック信号に応じて、前記マージ・イネーブル信号を記録し、 前記第2のマルチプレクサは、前記初期のロード・パルス信号が論理上のローになるのに応じて、前記ローカル・データ選択信号を前記第1のDタイプフリップフロップのデータ入力に再循環させ、前記マージ・イネーブル信号は、前記初期のロード・パルス信号が論理上のハイになるのに応じて前記第1のDタイプフリップフロップに結合され、 前記制御ロジックは、モード制御ロジックを更に有し、 前記モード制御ロジック は、 第3のマルチプレクサを含み、前記第3のマルチプレクサは、前記初期のロード・パルス信号に結合された第1のデータ入力を含み、 第2のDタイプフリップフロップを含み、前記第2のDタイプフリップフロップは、前記第3のマルチプレクサに結合され、前記第3のマルチプレクサの出力に結合されたデータ入力と、前記第1のクロック信号に結合されたクロック入力と、反転バス・モード信号に結合されたクリア入力と、前記バス・マルチプレクサの前記選択入力 及び前記第3のマルチプレクサの第2のデータ入力 に結合されたデータ出力 と、を 含み、前記第2のDタイプフリップフロップは、前記反転バス・モード信号、前記初期のロード・パルス信号、及び、前記第1のクロック信号に応じて、前記データ出力上の前記データバス選択信号を生成させ、 ORゲートを含み、前記ORゲートは、前記初期のロード・パルス信号に結合された第1の入力と、遅れたロード・パルス信号に結合された第2の入力と、を有し、前記ORゲートは、前記初期のロード・パルス信号と、前記遅れたロード・パルス信号を、論理的にOR演算し、 ANDゲートを含み、前記ANDゲートは、前記ORゲートの出力に結合された第1の入力と、バス・モード信号に結合された第2の入力と、前記第3のマルチプレクサの選択入力に結合された出力と、を有し、 インバータを含み、前記インバータは、バス・モード信号に結合された入力と、前記第2のDタイプフリップフロップの前記クリア入力が結合された出力、を有し、前記インバータは、バス・モード信号に応じて、前記反転バス・モード信号を生成し、 第4のマルチプレクサを含み、前記第4のマルチプレクサは、前記初期のロード・パルス信号に結合された第1のデータ入力、前記ORゲートの前記出力に結合する第2のデータ入力、前記バス・モード信号に結合する制御入力、及び、前記第1のPISOシフトレジスタのロード/シフトバーのバー入力に結合する多重化された出力を有し、 前記第3のマルチプレクサは、前記反転バス・モード信号に応じて、データバス選択信号を、前記第2のDタイプフリップフロップのデータ入力に再循環させ 前記第4のマルチプレクサは、前記初期のロード・パルス信号か、又は、前記初期のロード・パルス信号と前記遅れたロード・パルス信号との両方を、前記第1のPISOシフトレジスタのロード/シフトバーのバー入力に、選択的に結合させる 請求項 4 に記載の集積回路。
- 6前記集積回路は、バッファ集積回路であり、 前記ローカル・データバスは、12ビット長であり、 前記一以上のレーンのそれぞれの前記バス・マルチプレクサは、前記データバス選択信号に応じて、前記ローカル・データバスのより下位の6ビットを前記第1のPISOシフトレジスタに選択的に結合し、前記ローカル・データバスのより上位の6ビットを前記第1のPISOシフトレジスタに選択的に結合する 請求項 5 に記載の集積回路。
Independent claims6
144 paragraphs, as filed
Embodiments of the invention generally relate to memory, especially to merging data from a memory buffer onto a serial data channel.
In a memory circuit, there is generally a memory read latency, which is the time it takes to read the memory circuit for valid data. Memory write latency is generally also the period required to hold valid data in a memory circuit for writing data to memory. Memory read latency and memory write latency may sometimes be buffered from the processor by cache memory. However, sometimes the requested data cannot be found in the cache memory. In those cases, the processor may then need to read or write data to the memory circuit. In this way, each memory read latency or memory write latency may be implemented (experienced) by the processor. If the memory circuits are different, the memory read latency and the memory write latency do not have to match from one memory circuit to the next. In that case, the memory read latency and the memory write latency performed (experienced) by the processor will be different.
In advance, the memory modules are coupled to the mother or host printed circuit board, thereby being coupled in parallel with the parallel data bus so that the parallel data can be read from and written to the memory. A parallel data bus has parallel data bitlines that are synchronized together to transfer at least one data byte or word of data at a time. Parallel data bit lines are generally routed across a distance from one memory module socket on a printed circuit board (PCB) to the other. This creates a first parasitic capacitance load. As the memory module is coupled to the memory socket, an additional parasitic capacitance load is generated on the parallel data bitline of the parallel data bus. Since there may be many memory modules to be plugged in, the additional parasitic capacitive load can be significant and can deactivate high frequency memory circuits.
One memory module is generally addressed at once by an address on the address line. One addressed storage module generally writes data to a parallel data bus at once. Other memory modules generally have to wait for data written to the parallel data bus to avoid collisions.
Parallel data bitlines can speed up data flow, while parallel data buses in memory can slow down read and write access to data between memory circuits and processors.
In the following detailed description of each embodiment of the invention, a number of specific details will be provided to give a complete understanding of the invention. However, it will be apparent to those skilled in the art that each embodiment of the invention may be practiced without these particular details. In other cases, well-known methods, procedures, components, and circuits are not detailed so as not to unnecessarily obscure aspects of each embodiment of the invention.
Generally, in each embodiment of the present invention, data called northbound (upstream serial link) data merge (NBDM) that converts a part of data on a high-speed link with its own data at high speed. -Provide a merge function. That is, in each embodiment of the invention, part of the ingress serial data traffic (eg, "idle packets or frames") on the serial data link, where the ingress serial data traffic directs local data. The process of deciding whether to insert and retransmitting ingress data traffic with local data inserted into it (eg serial / parallel conversion, assembling to frames, and depacketizing / deinterleaving data). Is converted to its local data without the internal core logic doing it.
The input serial data must be pre-assembled into frames and received in order for the core logic to transmit the local data. The input / output (IO) interface of a memory module is received from another memory module on the serial data link or a memory controller and buffered without having input serial data processing to send local data. The input serial data stream bypassing the internal core logic of the integrated circuit may simply be retransmitted. This can reduce the data latency of the serial data stream. The portion of the retransmitted serial data stream is sometimes referred to as "feedthrough data" or "feedthrough data" (FTD).
Without transmitting any local data, the I / O interface typically bypasses the chip's core logic and retransmits the received serial data stream. The core logic of the buffer memory chip sends a merge request to the I / O interface along with the local data when it needs to send the local data. Since the core clock that generates the local data is tuned to the frame clock of the high-speed serial data link of each embodiment of the present invention during training, the input / output interface is idle packets or frames. Data can be merged immediately at the appropriate frame boundaries for conversion.
Earlier, it was intended that the received serial data would be assembled into a frame, received by the core logic, and then retransmitted to the output link. In this case, if the core logic has local data to send on the output link, then some incoming data will be replaced with its own data and the data will be repacketized and serialized on the output link. This introduces a data latency of at least 2 frames to the data. Each embodiment of the invention receives and analyzes input data during normal operation to transform idle packets during initial training to allow local data to be merged into the output link. Set the merge timing without. In each embodiment of the present invention, the buffer memory integrated circuit can reduce the data latency of at least two frames of data down by a few bit intervals.
In one embodiment of the invention, an integrated circuit is provided that includes a serial I / O interface with at least one lane. Each lane of the serial communication channel may include a first parallel input serial output (PISO) shift register, a first multiplexer, and a serial transmitter combined together.
The first parallel input serial output (PISO) shift register has a parallel input coupled to the local data bus, a clock input coupled to the first clock signal, and a load / shift bar input coupled to the load signal. The first PISO shift register serializes parallel data on the local data bus to serialized local data on the first serial output.
The first multiplexer combines a first data input coupled to a first serial output, a second data input to receive feedthrough data, and a first selection coupled to a local data selection signal. Has a control input. The multiplexer merges the selectively serialized local data and the feedthrough data according to the local data selection signal into a serial data stream of the multiplexed output.
The serial transmitter has an input that couples with the multiplexed output of the multiplexer to receive the serial data stream. The serial transmitter drives the serial data stream over the serial data link.
The feedthrough data may be 2 bits long, while the parallel input to the PISO shift register may be 6 bits long and the serial output of the PISO shift register may be 2 bits long. In this case, the first multiplexer is such a multiplexing such that the serial transmitter receives two bit serial data streams and serializes them as a single bit serial data stream on the serial data link. It may be a 2-bit bus multiplexer in which the serial data stream of the converted output is 2 bits long.
Each lane has a first input for receiving resynchronized data, a second input for receiving resampled data, and a second input that couples to a local clock mode signal. May be further provided with a multiplexer. The second multiplexer selects between resynchronized data and resampled data as feedthrough data according to the output local clock mode signal. Each lane may further include control logic that couples with a first multiplexer and a first PISO shift register. The control logic may include merge control logic and mode control logic. The control logic merges the serialized local and feedthrough data into a serial data stream in response to the first clock signal and the merge enable signal and the first clock signal. A merge enable signal for generating a local data selection signal may be received.
In another embodiment of the invention, receiving an input serial data stream representing feed-through frames of data interspersed between idle frames of data, merging without decoding the input serial data stream. Depending on the enable signal, both local frames of data and feedthrough frames of data are merged into an output serial data stream and output on the northbound data output to the next memory module or memory controller. Methods are provided for memory modules that include sending serial data streams. Local frames of data can be merged into the output serial data stream by converting idle frames of data in the input serial data stream. In the receiving input serial data stream, sampling (also referred to as resampling) the bits of the data in the input serial data stream, or the data in the input serial data stream. Resynchronization of bits of may be provided. In maraging both the local frame of the data and the data in the feedthrough frame, serializing the parallel bits of the local frame of the data into the serial bits of the data, and depending on the merge enable signal, of the data It may be provided to multiplex the serial bits of the data in the local frame with the serial bits of the feedthrough frame of the data into the serial bits of the output serial data stream. Local frames of data may be selectively received in parallel, in parallel, in 6-bit or 12-bit packets on the local bus, depending on the bus mode signal.
The system of another embodiment of the present invention is provided with a processor, a memory controller coupled to the processor, and at least one bank of memory coupled to the memory controller. Processors are provided to execute instructions and process data. The memory controller is provided to receive a write memory instruction having write data from the processor and to receive a read memory instruction from the processor and output the read data to it.
One bank of memory includes at least one memory module, each having a buffer integrated circuit coupled together and a random access memory integrated circuit. The buffer integrated circuit reads to a southbound (downstream serial link) serial input / output interface having at least one serial lane and at least one serial lane and memory controller of the northbound serial input to receive write data from the memory controller. Includes a northbound serial input / output interface with a northbound serial output for transmitting the data.
Each serial lane of the northbound I / O interface has a parallel / serial converter and a first multiplexer. The parallel / serial converter has a parallel bit on the local data bus and a parallel input coupled to a clock input coupled to the first clock signal and a load / shift-bar input coupled to the load signal. The parallel / serial converter serializes the parallel bits of data on the local data bus into serialized local data on the first serial output. The first multiplexer is a first data input that couples to the serial output of a parallel / serial converter, a second data input to receive serial feedthrough data from the northbound serial input, and a local data selection. It has a selective input that couples to the signal. A first multiplexer that selectively merges serialized local data and serial feedthrough data into a serial data stream on a northbound serial output in response to a local data selection signal.
Each serial lane on the northbound serial I / O interface provides an input that combines the serial data stream on the northbound serial data output with the multiplexed output of a first multiplexer that receives the serial data stream. You may have more transmitters to drive towards the memory controller you have (transmitters).
Each serial lane of the northbound serial input / output interface may further include control logic coupled to a multiplexer and a first parallel / serial converter. The control logic receives the first clock signal and the merge enable signal for generating the local data selection signal, and serializes the local data according to the merge enable signal and the first clock signal. And serial feedthrough data into a serial data stream.
For each bank of system memory, the memory controller has a northbound serial input interface that receives at least one lane of serial data from at least one memory module, and at least one of the serial data in at least one memory module. It has a southbound serial output interface that transmits two lanes.
In another embodiment of the invention, the buffered memory module is provided with a printed circuit board, a plurality of random access memory (RAM) integrated circuits, and a buffer integrated circuit. The printed circuit board has an edge coupling that couples to the receptacle of the host system. The plurality of random access memories are coupled to the (RAM) integrated circuit, and the buffer integrated circuit is coupled to the printed circuit board. The buffer integrated circuit is electrically coupled to a plurality of RAM integrated circuits and edge coupling. The buffer integrated circuit comprises a southbound I / O interface with data merge logic having multiple merge logic slices for multiple lanes of the serial data stream, and a northbound input / output interface.
Each merge logic slice of the buffer integrated circuit contains a first parallel input serial output (PISO) shift register and a first multiplexer. First Parallel Input-The serially output (PISO) shift register is the local data bus, the clock input coupled to the first clock signal, and the parallel input coupled to the load / shift bar input coupled to the first load signal. Has. The first PISO shift register serializes parallel data on the local data bus to serialized local data on the first serial output. The first multiplexer has a first data input coupled to the first serial output of the first PISO shift register, a second data input to receive serialized feedthrough data, and a local. It has a first selection input that couples to the data selection signal. The first multiplexer selectively merges the serialized local data with the serialized feedthrough data in response to the local data selection signal to serialize the serial data on the multiplexed output. Make it a stream.
Each merge logic slice may further include a first multiplexer and control logic that couples to a first PISO shift register. The control logic merges the merge enable signal and the serialized local data and the serialized feedthrough data into a serial data stream according to the first clock signal. The first clock signal and the merge enable signal are received to generate the selection signal.
The northbound I / O interface of the buffer integrated circuit of the buffered memory module provides multiple transmitters, serial data streams, to each merge logic slice, each with an input coupled to the corresponding output of the first multiplexer. It may further be equipped with multiple transmitters to receive and drive it over a serial data link.
In another embodiment of the invention, the memory system is provided with a plurality of buffered memory modules both daisy-chained together to form a bank of memory. Each buffered memory module includes a plurality of memory integrated circuits and a buffer integrated circuit coupled to the plurality of memory integrated circuits. The buffer integrated circuit is a southbound input / output serial interface that receives southbound serial data from a memory controller or a conventional buffered memory module and retransmits it to the next buffered memory module, serialized feedthrough. Northbound I / O serial interface that receives northbound serial data from at least one buffered memory module as data and retransmits it to the memory controller, addressed to the buffered memory module by a write command Write data that stores write data from the southbound input / output serial interface Write data first-in first-out method (FIFO) buffer, write data Transfer the write data stored in the FIFO buffer to at least one of multiple memory integrated circuits to transfer multiple read data. Read data from at least one of the integrated circuits of the memory input / output interface to transfer to the FIFO buffer, and from at least one of the integrated circuits as local data addressed from the buffered memory module by the read command. Contains a read data FIFO buffer, which stores read data.
The northbound I / O serial interface serializes local data from multiple memory integrated circuits, merges it, and serializes the received northbound serial data on a timing basis without decoding it. Make a northbound serial data stream with fed-through data. The northbound I / O serial interface comprises a third FIFO buffer, data merging logic that couples to the third FIFO buffer, and multiple transmitters that couple to the data merging logic.
The data merge logic is a first parallel input serial output (PISO) shift register, each serializing parallel data on the local data bus into serialized local data on the first serial output. And, in response to the local data selection signal, the first, which selectively merges the serialized local data with the serialized feedthrough data into a serial data stream on the multiplexed output. Equipped with a multiplexer. The PISO shift register has a local data bus, a clock input coupled to a first clock signal, and a parallel input coupled to a load / shift bar input coupled to a first load signal. The first multiplexer is used for the first data input coupled to the first serial output of the first PISO shift register, the second data input for receiving serialized feedthrough data, and the local data selection signal. It has a first selection input, to combine.
Each of the plurality of transmitters has an input coupled to the output corresponding to the first multiplexer of each merge logic slice. Multiple transmitters receive data from a serial data stream and drive it over a serial data link.
In the memory system, each merge logic slice of the data merge logic receives the first multiplexer and the first clock signal and merge enable signal, serialized local data, and serialized feed. It may further include control logic coupled to a first PISO shift register to generate a local data selection signal that merges through data into a serial data stream.
The memory system may further include a memory controller that couples into at least one of a plurality of buffered memory modules. The memory controller has a southbound output serial interface that sends the southbound serial data stream to at least one of the buffered memory modules and a northbound serial data stream to at least one of the buffered memory modules. It has a northbound input serial interface that receives from one.
See FIG. 1A here. A block diagram of a representative computer system 100 in which each embodiment of the present invention may be utilized is shown. The computer system 100A includes a central processing unit (CPU) 101, an input / output device (I / O) 102 such as a keyboard, a modem, a printer, an external storage device, and a monitor device such as a CRT or a graphic display. M) 103, is provided. The monitoring device (M) 103 may provide computer information in a human-understandable format, such as a visual or audio format. System 100 may be many different electronic systems other than computer systems.
See FIG. 1B here. A client-server system 100B in which each embodiment of the present invention may be utilized is shown. The client-server system 100B includes at least one client 110A-110M coupled to the network 112 and a server 114 coupled to the network 112. Clients 110A-110M communicate with server 114 through network 112 to send or receive information to gain access to any database and / or application software that the server may need. The server 114 may have a central processing unit with memory and may further include at least one disk drive storage device. Server 114 may be used in a storage area network (SAN), for example as a network attached storage (NAS) device, or may have a disk array. Data access to server 114 is shared through network 112 with multiple clients 110A-110C.
Here, by referring to FIG. 2A, a block diagram of a central processing unit 101A in which each embodiment of the present invention may be utilized is shown. The central processing unit 101A includes a processor 201, a memory controller 202, and a first memory 204A of the first memory channel coupled together as shown. The central processing unit 101A may further include a memory controller 202, a cache memory 203 coupled between the processors 201, and a disk storage device 206 coupled to the processor 201. The central processing unit 101A may further include a second memory channel having a second memory 204B coupled to the memory controller 202. As shown in the central processing unit 101A, the memory controller 202 and the cache memory 203 may be outside the processor 201.
See FIG. 2B here. A block diagram of another central processing unit 101B in which each embodiment of the present invention may be utilized is shown. The central processing unit 101B includes an internal memory controller 202'and a processor 201'with a first memory channel associated with the memory 204A coupled to the internal memory controller 202'in the processor 201'. Processor 201'may further include internal cache memory 203'. The central processing unit 101B may further include a second memory 204B for the second memory channel and a disk storage device 206 coupled to the processor 201'.
The disk storage device 206 may be a floppy disk, a zip disk, a DVD disk, a hard disk, a rewritable optical disc, a flash memory, or another non-volatile storage device.
Processors 201, 201'may further include at least one execution unit and at least one level of cache memory. Other levels of cache memory may be external to the processor and at the interface of the memory controller. The processor, at least one execution unit, or at least one level of cache memory may read or write data (including instructions) through a memory controller with memory 204A-204B. When interfacing with memory controllers 202, 202', there may be address, data, control and clock signals coupled to memory as part of the memory interface. Processors 201, 201'and disk storage device 206 may both read and write information to memory 204A, 204B.
The memories 204A and 204B shown in FIGS. 2A-2B are, for example, dual inline (FB) memory modules with full buffer (DIMM), (FBDIMM), or single inline memory modules with full buffer (FB) (SIMM). ), (FBSIMM), at least one buffered memory module (MM1-MMn) may be provided.
The memory controllers 202 and 202'interface each memory 204A-204B. In one embodiment of the invention, the memory controllers 202, 202'are, in particular, the buffers in the first buffered memory module MM1 of each memory 204A-204B (not shown in FIGS. 2A-2B, but the buffer of FIG. 5). Interface 450A). Interfaceing the buffer of a memory module with memory controllers 202, 202'can avoid a direct interface to the memory device of the buffered memory module (MM1-MMn). In this way, different types of memory elements can be used to provide memory areas when the buffer and the interface between the memory controllers can be kept consistent.
See FIG. 3 here. A buffered memory module (BMM) memory controller (BMMMC) 302 coupled to at least one memory bank 304A-304F (commonly referred to as memory bank 304, or multiple memory banks 304) is illustrated. The memory controller 302 can support two or more channels of memory and two or more memory banks of memory modules. Each memory bank 304 is composed of a plurality of buffered memory modules 310A-310H coupled together in a serial chain. This serial chain of buffered memory modules 310A-310H is sometimes referred to as the daisy chain. Adjacent memory modules are shown to be coupled together and sometimes daisy-chained together, for example, memory module 310A coupled to adjacent memory modules 310B.
Each memory module 310A-310H in each bank communicates bidirectionally along the serial chain of the memory modules 310A-310H in a serial form with a memory controller 302. There is a Southbound Serial Data Link (SB) from the memory controller 302 to each memory bank 304 that may be referred to as an output data link with output commands (eg, read / write) and data. All write data that is to be written to the memory module from the memory controller is transmitted over the southbound serial data link. There is a northbound serial data link (Nb) from each memory bank 304 to the memory controller 302, which may be referred to as a return data link with return data. All read data from the memory module is sent to the memory controller on the northbound serial data link.
In the Southbound Serial Data Link (SB), the data output from the memory controller 302 to the first memory bank 304 can read the data and pass it through the memory module 310B. Combined with memory module 310A. Memory module 310B can read data and pass it through the next memory module in the serial chain, and so on until it reaches the last memory module in the southbound serial chain. Will be done. The last memory module in the Southbound serial chain, the memory module 310H, no longer has a memory module to pass more data through, so the Southbound serial data link ends.
On the northbound serial data link (Nb), data is serially communicated from memory bank 304 in the direction of memory controller 302. Each memory module in each memory bank communicates the controller on the northbound serial data link (Nb) towards memory in the return direction. The memory module 310H initiates a serial chain of memory modules that are passing data towards the memory controller. The serial data transmitted by the memory module 310H is passed or otherwise retransmitted by the memory module 310G.
When the memory module 310G may pass or retransmit the serial data from the previous memory module 310H, it also transfers its own local data to the northbound serial data stream towards the memory controller 302. You may add or merge. Similarly, each memory module going down the chain either passes or retransmits the serial data from the previous memory module, sending its own local data to the northbound serial data stream towards memory controller 302. Add or merge. The last memory module in the northbound serial chain, memory module 310A, sends the final northbound serial data stream to memory controller 302.
Northbound and Southbound serial data links may be considered to provide point-to-point communication from one memory module to another along the serial chain as well. .. The serial data flowing from the memory controller 302 to the memory module 310A to the memory module 310H may be referred to as a southbound data flow. The serial data flow from the memory module 310H through the memory module 310Z to the memory controller 302 may be referred to as northbound data flow. In FIG. 3, the southbound dataflow is indicated by the arrow labeled SB, while the northbound dataflow is illustrated by the arrow labeled NB.
Here, by referring to FIG. 4, a buffered memory module (BMM) 310 exemplifying the memory modules 310A-310H is illustrated. The buffered memory module 310 may be of any type, for example SIMM or DIMM. The buffered memory module 310 includes a buffer integrated circuit chip (buffer) 450 coupled to the printed circuit board 451 and a memory integrated circuit chip (memory element) 452. The printed circuit board 451 includes an edge connector for connecting the printed circuit board to the edge connector of the host, or an edge connection 454. The southbound data input (SBDI) and northbound data output (NBDO) of memory module 310 are received or received from the previous buffered memory module or buffered memory controller, respectively. Will be sent. The northbound data input (NBDI) and southbound data output (SBDO) of memory module 310 are received or transmitted from, if any, to the next buffered memory module, respectively. ..
See FIG. 4 here. The memory controller 302 communicates with the buffer 450 of each memory module 310A-310H in each memory bank 304 by using the southbound data flow and the northbound data flow. The edge connection 454 of the memory module 310A, which is the first memory module closest to the memory controller of each bank, couples the buffer 450 of each memory module 310A to the memory controller 302. Memory module 310A no longer has adjacent memory modules in the path of northbound data flow. Northbound data flow from memory module 310A is coupled to memory controller 302. Adjacent memory modules 310A-310H in each bank are combined together to allow data to be written, read, and passed through each buffer 450 in each memory module. The last memory module farthest from the memory controller in each bank, memory module 310H, no longer has adjacent memory modules in the path of the southbound data flow. In this way, the memory module 310H does not pass any further southbound data flow along the serial chain of memory modules.
The memory controller 302 does not directly couple with the memory device 452 in any memory module. The buffer 450 in each memory module 310A-310H in each memory bank 304 is directly coupled to the memory device 452 on the printed circuit board 351. The buffer 450 provides data buffering to all memory integrated circuit chips or devices 452 on the same printed circuit board 451 of the memory module 310. The buffer 450 further performs serial / parallel conversion and parallel / serial conversion of data as needed, as well as performing interleaving / deinterleaving and packetizing / depacketizing of data. The buffer 450 also controls a portion of the Northbound and Southbound datalink serial chains by adjacent memory modules. Further, in the case of the first memory module, the memory module 310A, the buffer 450 also controls a part of the northbound and southbound datalink serial chains by the memory controller 302. Further, in the case of the last memory module, the memory module 310H, the buffer 450 also initializes the serial chain of the memory module, generates idle packets of data in idle frames or northbound data links, It also controls the northbound data flow to the memory controller 302.
In the absence of a direct coupling between the memory controller 302 of the memory module and the memory device 452, the memory chip or device 452 may differ in type, speed, size, etc., depending on where the buffer 450 communicates. .. This is an improved memory chip in the memory module by purchasing a new host or printed circuit board on the motherboard without the need to update the hardware interface between the memory controller and the memory module. Allows to be used. Memory modules that connect the printed circuit board to the host or motherboard are updated instead. In one embodiment of the invention, the memory chip, integrated circuit, or device 452 is a DDR memory chip with dynamic random access memory (DRAM). Otherwise, in each of the other embodiments of the invention, the memory chip, integrated circuit, or device 452 may be any other type of memory, or storage device.
See FIG. 5 here. One memory bank 304 coupled to the buffered memory module (BMM) memory controller 302 of the memory system memory banks 304A-304F is illustrated in more detail. In one embodiment of the invention, the BMM memory controller 302 is a full buffered dual inline (FBD) memory controller, and each memory module 310A-310H is a full buffered dual inline (FBD) memory module (FBDIMM). Is. The memory bank 304 includes at least one memory module 310A-310n that is daisy-chained together. Each memory module 310 acts like a repeater for valid data flowing in a serial bitstream along a northbound data link (Nb) and a southbound data link (SB).
Each memory module 310A-310n of the memory bank 304 includes a buffer 450A-450n, respectively. Each buffered memory module 310A-310N comprises a memory device 452A-452N which may be different from each other. For example, the memory device 452A in the buffered memory module 310A may be different from the memory device 452B in the buffered memory module 310B. That is, the buffer 450 of each memory module creates the type of memory used for the transparent memory element from the memory controller 302.
The buffer 450 of each memory module acts like a repeater for data flowing in a serial bitstream along a northbound data link (Nb) and a southbound data link (SB). In addition, each memory module's buffer 450 draws its own local data along the northbound data link (Nb) instead of idle or invalid frames of data, or partial frames. It may be inserted or merged into the lane of the flowing serial bitstream.
In order to synchronize the timings of the memory controller 302 and the memory modules 310A-310n in the memory bank 304 together, each memory module and a clock generator 500 coupled to the memory controller are provided. The clock signal 501 from the clock generator 500 is coupled to the memory controller 302. The clock signals 502A-502n are coupled to buffers 450A-450n of the memory modules 310A-310n, respectively.
The memory controller 302 communicates over the southbound data link SBl-SBn through a memory module in the memory bank 304. The memory controller 302 may receive data from each memory module 310 in the memory bank 304 over the northbound data link NB1-NBn. The Southbound Datalink SBl-SBn may include at least one lane of serial data. Similarly, the northbound data link NBl-NBn may include at least one lane of serial data. In one embodiment of the invention, the serial data of the northbound data link NBl-NBn has 14 lanes.
The last memory module, the memory module 310n, generates a bitstream of pseudo-random numbers, whether or not it has data to send, and directs it towards the memory controller 302 on the northbound link NBn. Start flowing to. A bitstream of pseudo-random numbers may be passed from one memory module on the northbound link NB1-NBn to the next. If memory module 310n has local data to send to memory controller 302, memory module 310n generates a frame of data with local data and uses that frame instead of a frame of data in a pseudorandom bitstream. , Placed on Northbound Link NBn. A bitstream of pseudo-random numbers may include a sequence of bits packetized into frames of data that indicate idle frames of data. Data idle frames are further completely replaced with other memory modules (memory modules 310A-310n) to merge frames of local data into a serial bitstream flowing over the northbound link NB1-NBn. May be done. For example, memory module 310B may receive idle frames on the input northbound link NB3, and instead of idle frames, it may receive frames of local data and a serial bitstream on the output northbound link NB2. May be merged into.
The memory system illustrated in FIG. 5 may further include an SM bus (SMBus) 506 coupled from the memory control 302 to each memory module 310A-310N. The SM bus 506 may be a serial data bus. The SM bus 506 is a sideband mechanism for accessing the internal registers of the buffer. The link parameters may be set by the buffer's BIOS before raising the northbound and southbound serial data links. The SM-bus may also be used to debug the system by accessing the buffer's internal registers.
The memory controller 302 may be part of a processor (as illustrated in processor 201'and memory controller 202' in FIG. 2B) or is illustrated in processor 201 and memory controller 202 in FIG. 2A. It may be a separate integrated circuit. In any case, the memory controller 302 receives a write memory instruction with write data from the processor for each write or read data to or from memory, receives a memory instruction read from the processor, and receives the processor. Read data can be output to. The memory controller 302 may include a Southbound Serial Output Interface (SBO) 510 for transmitting at least one lane of serial data to at least one memory module in each bank of memory. The memory controller 302 may further include a Northbound Serial Input Interface (NBI) 511 for receiving at least one lane of serial data from at least one memory module in each bank of memory.
Now refer to Figure 6 (Figures 6-1 and 6-2). A functional block diagram of buffer 450 of the buffered memory module 310 is shown. The buffer 450 is an integrated circuit that can be mounted on the printed circuit board 451 of the buffered memory module 310. The buffer 450 includes a southbound buffer input / output interface 600A and a northbound buffer input / output interface 600B for binding data to and from the buffered memory module 310.
The northbound buffer input / output interface 600B interfaces the northbound data output (NBDO) 601 and the northbound data input (NBDI) 602. The southbound buffer input / output interface 600A interfaces the southbound data input (SBDI) 603 and the southbound data output (SBDO) 604. In one embodiment of the invention, the northbound data input 602 and the northbound data output 601 include 14 lanes of a serial data stream. In one embodiment of the invention, the southbound data input 603 and the southbound data output 604 include 10 lanes of a serial data stream.
To interface with memory device 452, buffer 450 includes memory I / O interface 612. At memory I / O interface 612, DRAM data is transported in both directions over the DRAM data / strobe bus 605, while addresses and commands are sent to the memory device on the DRAM address / command bus 606A-606B. Transported on. The memory element 452 is timed by the DRAM clock bus 607A-607B to synchronize the data transfer with the memory input / output interface 612. From the core logic of buffer 450, memory I / O interface 612 commands from the multiplexer 635 on the CMD out bus 692, the address on the ADD out bus 693 from the multiplexer 637, and the data out bus from the multiplexer 636. Receives write data, on 691. The write data on the data out bus 691 is communicated to the corresponding memory device on the DRAM data / strobe bus 605. Address data on the data out bus 691 is communicated to the corresponding memory device on the DRAM address / command bus 606A-606B. Commands on the CMD outbus 692 are communicated to the appropriate memory device on the DRAM address / command bus 606A-606B.
The coupled phase-locked loop (PLL) 613 receives a reference clock (REF CLOCK) 502 to generate the core clock signal 611 for the functional block of buffer 450. The reference clock (REF CLOCK) 502 may be a differential input signal and is appropriately received by the differential input receiver. The buffer 450 further receives the SM bus 506 coupled to the SM bus controller 629. The reset signal (Reset #) 608 is coupled to the buffer 450 and the reset control block 628 to reset the functional block when it goes active low.
Between the memory I / O interface 612 and the buffer I / O interfaces 600A-600B is the core logic of the buffer 450. The core logic of buffer 450 is used to read data from memory devices and output data as local data through the northbound data interface 600B. In addition, any other response from the memory module is output by the buffer and entered into the northbound serial data stream through the northbound data interface 600B. The core logic of buffer 450 is also used for write data to the memory device received from the southbound data interface 600A. Commands for reading and writing data are received from Southbound Data Interface 600A. If the memory device 452 of the given buffered memory module 310 is inaccessible, serial data on northbound data input 602 and southbound data input 603 will pass through buffered I / O interfaces 600A-600B, respectively. , Northbound data output 601 and Southbound data output 604 may be passed. In this way, the data from the other buffered memory modules 310 passes through the memory controller on the northbound data interface 600B without being processed by the core logic of the buffer 450. Similarly, data from the memory controller may pass over other memory modules on the southbound data interface 600A without being processed by the core logic of buffer 450.
The core logic of buffer 450 has functional blocks for reading data from and writing to memory device 452. The core logic of buffer 450 is phase-locked loop (PLL) 613, data CRC generator 614, read FIFO buffer 633, 5-input 1-output bus multiplexer 616, synchronous and idle pattern generator 618, NB. LAI buffer 620, integrated built-in self-tester for link (IBIST) 622B, link initialization SM and control and configuration status register (CSR) 624B, reset controller 625, core control and configuration status register It includes (CSR) block 627, LAI controller block 628, SM bus controller 629, external MEMBIST memory calibration block 630, and failover block 646B coupled together as shown in FIG. The core logic of buffer 450 is command decoder and CRC checker block 626, idle built-in self-tester (IBIST) block 622A, link initialization SM and control and CSR block 624A, memory state controller and CSR632, write. Data FIFO buffer 634, 4-input 1-output bus multiplexer 635, 4-input 1-output bus multiplexer 636, 3-input 1-output bus multiplexer 637, LAI logic block 638, initialization pattern block 640, 1 bus multiplexer It has a failover block 646A with two to 642 and both combined as shown in Figure 6.
The multiplexer has at least two data inputs, an output, and at least one control input or selection input for selecting the data input given to the output of the multiplexer. For a two-input multiplexer, one control or select input is used to select the data that is the output of the multiplexer. The bus multiplexer receives a bit at each data input and also has an output containing the bit. A two-input, one-output bus multiplexer has its data input and two buses as a single bus output. A 3-input 1-output bus multiplexer has its data input and 3 buses as a single bus output. A 4-input 1-output bus multiplexer has its data input and 4 buses as a single bus output.
Within buffer 450, each buffer I / O interface 600A-600B has a FIFO buffer 651, a data merge logic 650, a transmitter 652, a receiver 654, a resynchronization block 653, and a demultiplexer / serial parallel converter block 656. Have. Data can pass through each buffer I / O interface 600A-600B through resynchronization path 661 or resampling path 662 without interfacing to the core logic. According to each embodiment of the present invention, the local data associated with the buffer 450 overwrites the idle frame without having the core logic to receive the serial data stream, in which the idle frame is contained. Can be merged into a serial data stream to determine where it is located.
The multiplexer 616 selects which data is directed to FIFO buffer 651 of northbound buffer I / O interface 600B for output as local data on the serial lane of northbound data output 601. In general, the multiplexer 616 has core control and status from CSR block 627 or read data from the read FIFO buffer 633, read data with attached CRC data from CRC generator 614, pattern generator 618. Synchronous or idle patterns from, or other control information from test pattern data from IBIST block 622B may be selected.
The multiplexer 642 selects which data is directed to the FIFO buffer 651 of the southbound buffer I / O interface 600A for output onto the serial lane of the southbound data output 604. In general, the multiplexer 642 may choose the initialization pattern from the initialization pattern block 640 or the test pattern data from the IBIST block 622A.
Now refer to FIG. 7A. A block diagram of the data merge logic 650 coupled to the transmitter 652 is shown. Transmitter 652 consists of N lanes of transmitter 752A-752n. As described above, in one embodiment of the present invention, the number of lanes is 10. In another embodiment of the invention, the number of lanes is 14. In the data merge logic 650, there are data merge logic slices 700A-700n for each of the N lanes.
The parallel local data bus 660 from the first-in first-out (FIFO) buffer 651 joins each data merge logic slice 700A-700n. Each lane of serial data in resynchronization path 661 joins each individual data merge logic slice 700A-700n. The bit length of the resynchronization bus 661 is twice the number of lanes. Two bits of each individual lane of resynchronization bus 661 are combined into individual data merge logic slices 700A-700N. Each lane of serial data on bus 662 to be resampled joins each individual data merge logic slice 700A-700n. The bit length of the resampling bus 662 is twice the number of lanes. The two bits of each individual lane of resampling bus 662 are combined into individual data merge logic slices 700A-700N.
Both the resampling bus 662 and the resynchronization bus 661 For each lane, a 2-bit serial data stream is transferred to each individual data merge logic slice 700A-700N. In contrast, the parallel data bus 660 combines 6 or 12 bits for each lane into individual data merge logic slices 700A-700N. The bit length of the parallel local data bus 660 is 12 times the number of lanes. But in 6-bit mode, only 6 of the 12 bits can be active per lane. The output from each data merge logic slice 700A-700N is a 2-bit serial data stream coupled to each serial transmitter 752A-752N. Each serial transmitter 752 converts two parallel bits of serial data into individual lanes 601A-601N or southbound data outputs (NBDO) 604, as shown in Figure 7A. Convert to a single bit serial data stream on individual lanes 604A-604N in SBDO) 601.
See FIG. 7B here. A schematic of the data merge logic slice 700i coupled to the transmitter 752i is shown. The data merge logic slice 700i represents one of the data merge logic slices 700A-700n for each N lane illustrated in Figure 7A. Transmitter 752i represents one of the transmitters 752A-752n for each N lane illustrated in Figure 7A.
Each data merge logic slice 700i is available in 12-bit full-frame mode (also known as 12-bit mode) or 6-bit long half-frame mode (also known as 6-bit mode). It can operate in one of the two-bit length modes of. The mode control signal (6bit_mode) 722 indicates and controls which 2-bit length mode core logic the data merge logic slice 700i operates in.
In full frame mode, or 12-bit mode, the core logic uses all 12-bit frames to communicate data on bus 660i using the data merge logic slice 700i. The lower 6 bits of bus 660i are represented by Data [5: 0] bus 726, while the upper 6 bits of bus 660i are represented by Delayed_data [5: 0] bus 727. The 12 bits of local data (Data [5: 0] and Delayed_data [5: 0]) merged into the serial data stream and transmitted are each subordinated by the "Early_Load_Pulse" control signal 720 at the beginning of the frame. It is latched by the parallel input serial output (PISO) converter 708B and the upper parallel input serial output (PISO) converter 708.
The lower parallel input serial output (PISO) converter 708B and the upper parallel input serial output (PISO) converter 708A are parallel input serial output (PISO) shift registers, which are also referred to herein. Will be done. Each PISO converter 708A-708B, also referred to as the PISO shift register 708A-708B, has parallel data inputs, clock inputs, load / shift bar inputs, serial inputs (SIN), and serial outputs (SO). The serial output of the upper PISO shift register 708A is coupled to the serial input of the lower PISO shift register 708B to support serializing the 12 parallel bits of the local data bus 660i. The serial input of the upper PISO shift register 708A is coupled to a logical low (eg ground) in one embodiment of the invention or a logical high (eg VDD) in another embodiment of the invention. Can be done. The serial output (SOUT) of the PISO shift registers 708A-708B is 2 bits at a time in one embodiment of the invention. In another embodiment of the invention, the serial output (SOUT) of the PISO shift registers 708A-708B may be 1 bit at a time.
In 12-bit mode, the 6 bits of bus 726 are coupled to the parallel data inputs (pins) of the lower PISO shift register 708B, while the 6 bits of bus 727 are the parallel data of the upper PISO shift register 708. Combined to the input (pin). These 12 bits are loaded into each PISO shift register during the initial load pulse 720 with a mode control signal 722 indicating 12-bit bus mode. (For example, in one embodiment of the invention, the mode control signal 722 indicates a 12-bit mode by being logically low level and a 6-bit mode by being logically high level. In 12-bit mode, the clear input to the D-type flip-flop 706A is logically high setting, and the Q output of the D-type flip-flop 706A is the control input to the multiplexer 703, bus 728. It is logically zero so that bus 726 is selected for output up.
In half-frame mode, or 6-bit mode, the core logic only uses 6-bit half-frames at a time to communicate data on the bus 660i with the data merge logic slice 700i. Is. The core logic sends 6 bits of data at a time, or initial data offset by half a frame (Data [5: 0] 726) and delayed data (Delayed_data [5: 0]). .. In half-frame mode, only the lower PISO shift register 708B of the data merge logic slice 700i is used to merge the data into the serial data stream for transmission.
The 6-bit mode multiplexer 703 selectively coupled 6 bits of bus 726 to the parallel data input (pin) of the lower PISO shift register 708B during the initial load pulse 720 and was delayed. During the load pulse 721, 6 bits of bus 727 are coupled to the parallel data input (pin) of the lower PISO shift register 708B. The six bits of bus 726 are loaded into the PISO shift register 708B by the mode control signal 722 indicating the 6-bit bus mode during the delayed load pulse 721. The 6 bits of bus 727 are loaded into the PISO shift register 708B by the mode control signal 722 indicating the 6-bit bus mode during the initial load pulse 720.
The data merge slice 700i has data path logic and control logic 701i. The datapath logic selectively merges local and feedthrough data into the serial bitstream. The control logic 701i controls the data path logic of each data merge slice in order to properly synchronize the merging of local data and feedthrough data to the serial bitstream.
The control logic 701i with mode control logic and merge control logic is shown in Figure 7B and, as illustrated, three single-bit 2-input 1-output multiplexers 702A-702C, set / It has a reset D flip-flop 706A-706B, an OR gate 710, an AND gate 711, and an inverter 712. The signal generated by the control logic 701i is coupled to the datapath logic. The multiplexer 702A-702B, D-type flip-flop 706A, OR gate 710, AND gate 711, and inverter 712 provide mode control logic. The multiplexer 702C and D-type flip-flop 706B provide merge control logic.
The datapath logic is a 6-bit 2-input 1-output bus multiplexer 703, a 2-bit 2-input 1-output bus multiplexer 704-705, a pair of 6-bit inputs, coupled together, as shown in Figure 7B. It has a parallel input serial output (PISO) converter 708A-708B with / 2 bit output.
Each slice 700i of data merge logic 650 receives a 2-bit serial lane of resynchronized data 661i, a 2-bit serial lane of resampling data 662i, and a 12-bit parallel lane of local data 660i. You may. The parallel lanes of the local data 660i are from the core logic of buffer 450 and may be of various types of data. For example, local data 660i was received, transmitted, or generated by read data from memory device 452, cyclic redundancy check (CRC) data, test data, status data, or the core logic of the buffer. It can be any other data.
The 2-bit lane of the resynchronized data 661i and the 2-bit lane of the resampled data 662i do not have a coupling point in the core logic of the given buffer 450 and are locally clocked by the multiplexer 705. Depending on the mode signal 736, it is multiplexed into feed-through data (also also referred to herein as "feed-through data") 725. When buffer 450 is operating in local clock mode, resynchronized data is multiplexed over feedthrough data 725. If buffer 450 is not operating in local clock mode, resampling data 662i is multiplexed over feedthrough data 725. In local clock mode, the phase-locked loop (PLL) clock generator sends a remote clock signal in the buffer used to resynchronize the input serial data stream to generate resynchronized data. Used to generate. When not in local clock mode, the received clock is generated from a frame of data in the received serial data stream that was used to sample the input serial data stream to generate resampling data. And be synchronized. The clock signal Clock_2UI723 is switched between a locally generated clock signal and a received clock signal according to the local clock mode signal 736. The source of feedthrough data 725 is from buffer 450 of another memory module 310 on the northbound (Nb) side (also referred to as transmitted northbound data), or from the southbound (SB) side (also referred to as). It may be from buffer 450 of another memory module 310 on (referred to as transmitted southbound data), and instead from memory controller 302 on the southbound (SB) side .
The 2-input 1-output bus multiplexer 704 has 2 bits of serial feedthrough data 725 as the first input, 2-bit serial output from the 6-2 PISO shift register 708B as the second input, and control. Receives the local data selection signal (PISO_SEL) 732 at the input. 6-2 The 2-bit serial output 735 from the PISO shift register 708B is the two serialized bits of the local data 735 from the parallel data bus 660i. Thus, depending on the local data selection signal (PISO_SEL) 732, the multiplexer 704 is generated by the 6-2 PISO shift register 708B from the 2-bit or parallel data bus 660i of the feedthrough data 725 and serialized. Both of the 2 bits of the local data 735 that have been created are selected for output. The 2-bit output 730 from the multiplexer 704 is coupled to the transmitter 752 and further serialized into a single bit over the lanes NBDOi / SBDOi601i, 604i. In this way, the local data from the core logic is multiplexed with the feedthrough data and can be merged into the lanes of the serial bitstream at NBDOi / SBDOi601i, 604i.
The local data selection signal (PISO_SEL) 732 that controls the merging of data to the serial bitstream is generated by the D flip-flop 706B. In response to the merge enable signal 724, the D flip-flop 706B generates a local data selection signal (PISO_SEL) 732 on the rising edge of the clock signal Clock_2UI723. The merge enable signal 724 is coupled to the first input of the multiplexer 702C. The local data selection signal (PISO_SEL) 732 is fed back and coupled to the second input of the multiplexer's 702C. The output of the multiplexer 702C is coupled to the D input of the D flip-flop 706B. The initial load pulse (EARLY_LD_PULSE) signal 720 is coupled to the selective control input of the multiplexer 702C. When the initial load pulse 720 is active high, the merge enable signal 724 is output by the multiplexer 702C and coupled to the D input of the D flip-flop 706B. When the initial load pulse 720 is low, the local data selection signal (PISO_SEL) 732 is fed back to the multiplexer 702C and of the D flip-flop 706B to hold the current state of the local data selection signal (PISO_SEL) 732. Combine to D input. The initial load pulse 720 is recorded periodically, so if the merge enable signal 724 is low, it clears the D flip-flop 706B at the appropriate time, so its Q output merges the data. Is a low logic level signal that terminates.
The merge enable signal 724 is synchronized with the local data selection signal (PISO_SEL) 732 on the edge of the clock signal Clock_2UI723. The merge_enable signal 724 is sampled during early_load_pulse720 to generate the local data selection signal (PISO_SEL) 732, and the multiplexer 704 is switched at the frame boundary (12 bits of data per frame lane). When the merge enable signal 724 is active high on the rising edge of the clock signal Clock_2UI723, the local data selection signal (PISO_SEL) 732 is a multiplexer 704 for selecting two serialized bits of the local data 735, It goes active high to control its 2-bit output 730, as well as. If the merge enable signal 724 is low on the rising edge of the clock signal Clock_2UI723, then the local data selection signal (PISO_SEL) 732 is its multiplexer 704, as well as the multiplexer 704, for selecting the two feedthrough bits of the data 725. It remains low to control the 2-bit output 730.
The two serial bits of the parallel data bus 660i are merged into lanes NBDOi / SBDOi601i, 604i., Depending on the local data selection signal (PISO-SEL) 732 being logically high. The two bits of feedthrough data 725 are selected by the multiplexer 704 for output on lanes NBDOi / SBDOi601i, 604i., Depending on the local data selection signal (PISO_SEL) 732 being logically low.
The local data selection signal (PISO_SEL) 732 responds to the merge enable signal 724, so the generation of the merge enable signal 724 is parallel to the bus 660i that is merged onto the serial data stream in lanes NBDOi / SBDOi601i, 604i. Give the data. The merge enable signal 724 is timely to provide local data to be merged into the serial data stream by the link control logic (link initialization SM and control and CSR functional block 624B shown in Figure 6). Is generated within the time of.
Temporarily return to Fig. 5 for reference. The timing of the merge enable signal is determined for each memory module 310 during system initialization and training. For the last memory module 310n in bank 304, the merge enable signal has more meaning than the data transmit signal, such as there are no more memory modules for the data being continuously generated on the northbound data link. Note that.
See FIG. 10 here. The flow chart is illustrated for initialization, training, and the behavior of the buffer when merging both local and feedthrough data into the serial data stream output. The flowchart starts at block 1000.
At block 1002, the buffer of each memory module in each memory bank is initialized. During the initialization of memory bank 304, each memory module initializes its southbound and northbound serial data links (also referred to as being part of link training). .. The memory controller 302 transmits an external initialization pattern on the Southbound (SB) data link SBl-SBn. During initialization, the last memory module 310n buffer 450n receives the initialization pattern on the southbound datalink SBn and passes it through the other memory modules to the northbound (Nb) datalink NB1-NBn. Retransmit the top back to memory controller 302. The initialization pattern received by the buffer on the northbound (Nb) data link NB1-NBn synchronizes the frames of each lane of serial data to lock the bits so that each buffer has its own clock. Used for purposes. The buffer clock may be synchronized with the initialization pattern. The timing of the logic is to receive a packet of data in the serial data stream, as well as to analyze headers from frames of data, and any error correction / detection, or other data fields in the packet. , May be matched to the initialization pattern. The generation of Early_Ld_Pulse720 is set to coincide with the beginning of a frame of data received by a given memory module. The generation of Late_LD_Pulse721 is set to be at the half-frame boundary of the frame of the data received by the given memory module.
Then in block 1004, each buffer of each memory module in each memory bank is trained. After sending the initialization pattern, the memory controller 302 sends the training pattern to the last memory module in bank 304 given during training. During training, the last memory module 310n buffer 450n is northbound through other memory modules to receive the training pattern on the southbound datalink SBn and return it to memory controller 302. Nb) Retransmit over datalink NB1-NBn. Each memory module observes one of the training patterns on the Southbound (SB) datalink and determines the time or clock cycle to return to the same memory module on the Northbound (Nb) datalink. decide. The round trip time is determined for a given position in each memory module.
Given that the requests are not over-concentrated, the round trip time is a given memory module for merging data onto the northbound data link without colliding with the valid data of other memory modules. Represents a slot in time that is safe for. With a given memory module, idle data packets are required to be received on the northbound datalink at the moment after checking the memory request command on the southbound datalink. .. At that point, the idle data packet can be replaced with a local data packet. The round trip time for the data delay time for a given memory module, and the command sets the timing of the merge enable signal used to control merging local data into the northbound data link. It is the basis for. If the round trip time is long, the data can be fetched first, placed in the FIFO buffer, and wait for a reasonable amount of time to be merged into the northbound data stream. The distance between the read / write FIFO buffer pointers of the buffer's northbound interface can be set based on the round trip timing.
The round trip time may be determined as a function of an integer of the period of the bit rate clock (clock_2UI 723). The number of memory modules in a channel and the command for the data delay of the last memory module in the channel determine the round trip time for that channel.
The command for data delay for each memory module may be further determined to assist in establishing the timing of the merge enable signal for each memory module. The command for data delay timing may have at least one of the following periods: Time for commands to be transferred from Southbound I / O interface 600A to memory I / O interface 612, time for commands to be transferred from memory I / O interface 612 to memory device 452, memory I / O interface 612, and memory Differences in clock timing of device 452, delays by route the clock signal to memory device 452, and command signals, buffer 450, and any setup / hold time for memory device 452, memory device 452 (eg, memory device 452). Read latency (for CAS timing and any additional latency), delay by route the data signal from memory device 452 to buffer 450, and strobe signal, delay skew of data between memory devices, memory I / O interface. Delay by 612, buffer 450, and any setup / hold time for memory device 452, and time for data to be transferred from memory I / O interface 612 to northbound I / O interface 600B (this is buffer 450). It may include a delay for the data to be buffered and timed in). The command for data delay timing depends on the number of multiple frames or the granularity of the delay time that exists as a function of an integer of the bit rate clock period (bit timing such as frame / 12 or clock_2ui / 2). It may be determined as a fraction of it. Commands for the data delay timing of memory modules, such as the last memory module 310n, can be increased by programmatically setting registers when additional delay times are required.
In the next block 1006, after initialization and training, each buffer is ready to receive the input serial data stream from the serial data input. However, the buffer of the last memory module 310n in memory bank 304 sends both idle packets and read-requested data packets towards memory controller 302 on the northbound data link. On the other hand, the input serial data stream receives data feedthrough frames scattered between idle frames of data.
In the next block 1008, the determination may be made with respect to the availability of local data. If there is local data to merge into the serial data stream, then the control flow jumps to block 1010. If there is no more local data to merge into the serial data stream, then the control flow jumps to block 1014.
At block 1014 with no local data to merge, feedthrough data is sent over the serial data output. The feedthrough data may have that bit of data in the resampled input serial data stream. Alternatively, the feedthrough data may have that bit of data in the resynchronized input serial data stream. The control flow then jumps back to block 1006 to continuously receive the input serial data stream.
At block 1010 for merging local data, frames of local data are replaced with feedthrough data in the output serial data stream. That is, if local data needs to be sent by a buffer, frames of data in the input serial data stream will be dropped and frames of local data will be sent at that location in response to the merge enable signal. Frames of local data and feedthrough data are merged together by serializing the parallel bits of the local frame of the data to the serial bits of the data, and then, depending on the merge enable signal, of the data. The serial bits of the data in the local frame and the serial bits of the feedthrough frame of the data may be multiplexed over the serial bits of the output serial data stream. During initialization and training, the host and memory controller ensure that idle frames of data in the input serial data stream are replaced with local frames of data. The buffer does not need to check if the input frame of the exchanged input serial data stream is a frame in which the data is idle.
At block 1012, the output serial data stream with the merged data is sent over the chain onto the serial data output to the next memory module, or instead to the memory controller.
The control process then jumps to block 1006 to continue receiving the input serial data stream from the serial data input.
As mentioned above, the core logic and local data from buffer 450 may be output of 6-bit chunks or 12-bit chunks at a time. The mode control signal (6bit_mode) 722 determines whether the data merge logic slice 700i operates in 6-bit mode (half-frame mode) or 12-bit mode (all-frame mode). .. The mode control signal (6bit_mode) 722 is coupled to the multiplexer 702A, the first input of the AND gate 711, and the selective input or control input of the input to the inverter 712.
The initial load pulse signal 720 controls the loading of the first 6 bits on the parallel data bus 660i. The delayed load pulse signal 721 controls the loading of the second 6 bits on the parallel data bus 660i. The delayed load pulse 721 is coupled to the first input of the OR gate 710. The initial load pulse control signal 720 was the first input of the multiplexer 702B, the second input of the OR gate 710, the first input of the multiplexer signal 702A, the load / shift-bar input of the 6-2 PISO shift register 708A, And it is coupled to the 702C select input of the multiplexer.
The clock signal Clock_2UI723 is coupled to the clock input of the D flip-FLOPS 706A-706B and the clock input of the 6-2 PISO shift register 708A-708B. The output of the multiplexer 702A is coupled to the load / shift-bar input of the 6-2 PISO shift register 708B.
6-2 The parallel input of PISO shift register 708A is coupled to the 6-bit delayed data bus 727. The 2-bit serial output of the 6-2PISO shift register 708A is coupled to the 2-bit serial input of the 6-2PISO shift register 708B. 6-2 The parallel input of the PISO shift register 708B is coupled to the 6-bit output from the multiplexer 703. Thus, when the data merge logical slice 700i is in 12-bit mode, 12-bit data can be loaded into the 6-2 PISO shift register 708A-708B and then 2 bits by the multiplexer 704. It is serially shifted from the serial output 708B and combined with the transmitter 752i.
The serial transmitter 752i is double clocked by a clock signal to set two parallel bits to one serial bit at its outputs 601i, 604i.
The data merge logical slice 700i is in 12-bit mode when the 6-bit_mode control signal 722 is logically low. The data merge logical / 700i is in 6-bit mode when the 6-bit mode control signal 722 is logically high. The control logic 710-712 with the multiplexer 702B and the D flip-flop 706A is coupled to the selective input of the multiplexer 703 to determine the 12-bit or 6-bit mode depending on the 6-bit mode control signal 722. Generates the data bus selection (Data_Sel) signal 729. When the data bus selection signal 729 is logically low, 12 bits of data are loaded in parallel with the 6-2 PISO shift registers 708A-708B. When the data bus selection signal 729 is logically high, the 6 bits of the data bus 727 are coupled to the 6-2 PISO shift register 708B.
In 6-bit mode, either the initial load pulse signal 720 or the delayed load pulse 721 can load parallel data into the 6-2 PISO shift register 708B. In either 6-bit mode or 12-bit mode, the initial load pulse 720 is used only to load parallel data from the data bus 727 into the 6-2 PISO shift register 708A.
6-2 The serial input of PISO shift register 708A is coupled to ground so that only zeros are serially shifted after the transmitted data. Instead, the serial input of 6-2PISO shift register 708A is coupled to VDD so that only the logical ones are serially shifted after the transmitted data.
The Q output of the D flip-flop 706A is coupled to the second input of the multiplexer 702B when the output of the AND gate 711 is logically low, and the Q output selects the loaded logical state on the data bus (DATA_SEL). ) Coupled to the D input of the D flip-flop 706A to hold in it of signal 729.
Now refer to FIG. A timing diagram of the waveform representing the data merge logic slice 700i operating in 12-bit mode is shown. That is, the 6-bit mode control signal 722 is a logical row in the timing diagram of FIG.
In FIG. 8, the Clock_2UI signal 723 is represented by the waveform 823. The core clock signal 611 is represented by waveform 811. The lower 6 bits of the data on the parallel data bus 690 (MEM_DATA input [5: 0]) 690A are indicated by the waveform 890A. The upper 6 bits of the data on the parallel data bus 690 (MEM_DATA input [11: 6]) 690B are shown in waveform diagram 890B. The lower 6 bits of the data (FBD_DATA [5: 0]) 726 on the parallel data bus 660i are shown in waveform diagram 826. The upper 6 bits of the data (FBD_DATA [11: 6]) 727 on the parallel data bus 660i are shown in waveform diagram 827. The merge enable control signal 724 is shown in waveform diagram 824. The initial load pulse control signal 720 is represented by waveform 820. The delayed load pulse control signal 721 is indicated by waveform 821. The local data selection control signal (PISO_SEL) 732 is illustrated in waveform 832. The single-bit serial output data stream NBDOi601i is represented by waveform 801.
Without any local data to merge into the northbound serial data stream, buffer 450 bypasses the bits received at northbound data input 602 (feedthrough data 725) through buffer 450's core logic. This allows it to pass through the transmitter 752i in the high speed clock area. The local data selection control signal (PISO_SEL) 732 is low when the feedthrough data 725 is multiplexed on the transmitter 752i, as illustrated in Waveform 832.
As mentioned above, the "Early_Ld_Pulse" 720 is set to coincide with the beginning of the frame (as referenced in the link) and late_ld_pulse721 is half-during the first training of the lane of the serial data link. Set to be at the frame boundary. In one embodiment of the invention, a frame of data is a logical unit of data on the link in the all-frame operation mode and is made up of 12-bit data.
In all frame operation mode, the 12 bits of the frame are loaded into the PISO shift register using the "Early_Ld_Pulse" signal 720. The "late_ld_pulse" signal 721 is not used to load the bits into the PISO shift register. Both the upper and lower PISO shift registers 708A-708B are used in this mode. The 6-bit_mode control signal 722 is low in 12-bit mode and is low in 12-bit mode by clearing the output of the D flip-flop 706A, resulting in a Data_Sel signal 729. When the "Data_Sel" signal 729 is low in 12 bit modes, the 6 lower data bits (FBD_DATA [5: 0] 726 on bus 660i) are moved by the multiplexer 703 to the lower PISO shift register. Combined with 708B.
The periodic generation of Early_Ld_Pulse720 also allows sampling of the "Merge_enable" signal 724 by the D flip-flop 706B. The periodic generation of Early_Ld_Pulse720 is active, high, and selectively controls the multiplexer 702C to select the merge_enable signal 724 as the output of the data coupled to the data input D of the D flip-flop 706B.
As mentioned above, the merge enable signal 724 is used to insert local data from a given memory module into the northbound serial data lane to exchange packets or idle frames of data in the serial data stream. In addition, it occurs at the right time. Waveform 824 is when local data is valid on bits above the data bus 660i (FBDJD ATA [11: 6]) 727 and below it (FBD_DATA [5: 0]) 726. , Indicates the active high pulse 844 generated.
When the active high pulse 844 occurs in the waveform 824 of the merge enable signal 724, the pulse 840A-840B of the early_ld_pulse signal 720 uses the clock_2UI signal 723 of the active high pulse 844 of the merge enable signal 724. Allow samples to be taken by flip-flop 706B. This causes an active high pulse 842 in the waveform 832 of the local data selection signal (PISO_SEL) 732. The active high pulse 842 of the local data selection signal (PISO_SEL) 732 gives 2-bit "feedthrough data" 725 at its output, and instead gives 2-bit serialized local data at its output. Switch the multiplexer 704 to give 735. Switching from feedthrough data 725 to local data 735 occurs at the frame boundary when the active high pulse 842 first occurs. This is because the falling edge of the "Early_Ld_Pulse" 720, which begins to shift to the PISO shift registers 708A-708B, coincides with the frame start position.
When both the merging data with the "Early_Ld_Pulse" 720 and the multiplexer output 731 are low, the PISO shift register 708A-708B serializes 12 bits of local data using the "Clock_2ui" clock signal 723. Two bits are serially shifted and output on the output 735 at a time. The transmitter 725i further serializes the two bits on the NBDoi output 601i into a single-bit serial data stream, as shown by the local data shown in waveform 801 above.
See FIG. 9 here. A timing diagram of the waveform representing the data merge logic slice 700i operating in 6-bit mode is shown. That is, the 6-bit mode control signal (6BITJVIODE) 722 is logically high as shown by the waveform 922 in the timing diagram of FIG.
In FIG. 9, the Clock_2UI signal 723 is illustrated with waveform 923. The core clock signal (core_clk) 611 is illustrated with waveform 901. The lower six parallel data bits (MEM_DATA input [5: 0]) 690A on the memory data bus 690 are represented by the waveform 990A. The upper six parallel data bits (MEM-data input [11: 6]) 690B of the memory data bus 690 are represented by the waveform 990B. The lower 6 bits of the data (FBD_DATA [5: 0]) 726 on the parallel data bus 660i are shown in waveform diagram 926. The upper 6 bits of the data (FBD_DATA [11: 6]) 727 on the parallel data bus 660i are shown in waveform diagram 927. The merge enable control signal 724 is shown by waveform diagram 924 and occurs earlier than that of waveform 824 in FIG. The initial load pulse control signal (EARLY_LD_PULSE) 720 is shown in waveform 920. The delayed load pulse control signal (LATE_LD_PULSE) 721 is indicated by waveform 921. The data bus selection control signal (DATA_SEL) 729 is indicated by the waveform 929. The local data selection control signal (PISO_SEL) 732 is shown in waveform 932. The single-bit serial output data stream NBDOi601i is represented by waveform 901.
In 6-bit mode, the lower PISO shift register 708B is used to convert the parallel bits of data into serial data with a shifted bit output. The data bus selection signal (DATA_SEL) 729 switches between the lowest 6 bits of the frame, FBD_Data [5: 0] 726, or the most significant 6 bits of the frame, FBD_Data [11: 6] 727, of the bus multiplexer 703. The selected output loads it into the lower PISO shift register 708B.
Both the "Early_Ld_Pulse" 720 and the "Late_Ld_Pulse" 721 were coupled by the multiplexer 702A to the load / shift bar input of the lower PISO shift register 708B by the multiplexer 702A when the 6BIT_MODE signal 722 was active high. Therefore, both load data and shift data can be output to the lower PISO shift register 708B.
When "Early_Ld_Pulse" 720 and "Late_Ld_Pulse" 721 are low, the bits are shifted and output from the lower PISO shift register 708B. Also, when the load / shift bar control input is high while the parallel load of bits is input to the lower PISO shift register 708B, the preloaded bits continue to be shifted and output. When the load / shift bar control input returns to low after parallel loading of data bits, the newly loaded bits are then shifted out by the lower PISO shift register 708B. Thus, all 6 bits of data may be shifted out when a new set of parallel bits is loaded.
The lowest 6 bits of the frame, FBD_Data [5: 0] 726, is the input waveform of "Early_Ld_Pulse" 720 when the data bus selection signal (DATA_SEL) 729 is low, for example low points 949C, 949D. The 920 pulses 940A and 940B load the lower PISO shift register 708B. The top 6 bits of the frame, FBD_Data [11: 6] 727, is the input waveform of "Late_Ld_Pulse" 721 when the data bus selection signal (DATA_SEL) 729 is high, for example between pulses 949A and 949B. It is loaded into the lower PISO shift register 708B by the pulse 941A and 941B of 921.
In 6-bit mode, switching between serialized "feedthrough data" 725 and serialized local data 735 is similar to the 12-bit mode behavior described above and is here for brevity reasons. I won't repeat it.
When merging data, the PISO shift register 708B outputs the most significant 6 bits or the least significant 6 bits of the local data that is being serially shifted and output 2 bits at a time using the Clock_2UI clock signal 723. Convert to the top of the serial output 735. The transmitter 725i further serializes the two bits on the NBDoi output 601i into a single-bit serial data stream, as illustrated by the local data shown in waveform 901 above.
During 6-bit mode, all frames of data are further transmitted and each embodiment of the invention is merged into a serial data stream, further reducing the latency of local data. When comparing FIGS. 8 and 9, the merging of local data is one frame earlier than in FIG.
Each embodiment of the present invention combines feedthrough data and local data with high speed serial without the decoding input packet of the serial input data stream to determine the location of the idle packet. -Allows you to marge on data links. The input serial data stream is pre-received, depacketized / decoded, and reassembled into frames by the retransmitted pre-core logic. Each embodiment of the invention prevents the input serial data stream from being depacketized / decoded and reassembled into frames of data, and prevents encoding / packetization for retransmission. Each embodiment of the invention allows retransmission of the input serial data stream and merging of local data to the serial data stream without the core logic of the buffer integrated circuit. In a multi-memory module system, the serial communication channel can continue to operate even if the memory integrates circuits and one input of the daisy-chained memory modules does not work.
Each embodiment of the invention is designed to provide low latency memory access operations. This can give each bank more memory with more memory modules, without memory access latency that degrades system performance as the number of memory modules in the channel increases.
An exemplary embodiment is described and shown in the accompanying drawings, but such embodiments are merely illustrated and do not limit the invention in a broad sense, and the present invention is illustrated. It is not limited to the specific structure described in the invention and its arrangement. This is because it can be understood by those skilled in the art that various other variants can be formed. For example, one embodiment of the invention is described to provide a serial data link for a dual inline memory module with a full buffer. However, each embodiment of the present invention may be implemented in other types of memory modules and systems. In another embodiment, the data is stored at once by merge logic on a 2-bit bus around the PISO shift registers 708A-708B to provide harmonious data timing in one embodiment of the invention. 2 bits are serialized. However, each embodiment of the invention uses single-bit output PISOs with different clock timings that serialize local data into feed-through data and single-bit serial data streams, with multiplexers 704 and 705 being single. May be provided to support bit-serial data streams.
<figref num="1A">A block diagram of a representative computer system in which each embodiment of the present invention may be used is shown.</figref>
<figref num="1B">A block diagram of a client-server system in which each embodiment of the present invention may be used is shown.</figref>
<figref num="2B">A block diagram of another central processing unit in which each embodiment of the present invention may be used is shown.</figref>
<figref num="3">A simplified block diagram of a buffered memory controller that combines input and output data into a bank of buffered memory modules is shown.</figref>
<figref num="4">A block diagram of a buffered memory module with a buffer in which data may be merged with feedthrough data is shown.</figref>
<figref num="5">A detailed block diagram of a buffered memory controller coupled to a bank of buffered memory modules is shown.</figref>
<figref num="6-1">The functional block diagram of the buffer of the memory module with a buffer is shown.</figref><figref num="6-2">The functional block diagram of the buffer of the memory module with a buffer is shown.</figref>
<figref num="7A">A simplified block diagram of the data merge logic including the lanes of the data merge logic slices coupled to the transmitter is shown.</figref>
<figref num="7B">A schematic of a data merge logic slice for one lane of serial data is shown.</figref>
<figref num="8">A timing diagram of the signal for a data merge logic slice operating in 12-bit mode is shown.</figref>
<figref num="9">A timing diagram of the signal for a data merge logic slice operating in 6-bit mode is shown.</figref>
<figref num="10">The flow chart of buffer initialization, training, and operation when merging both local data and feedthrough data into a serial data stream output is shown.</figref>
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO2004102403A2 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| WO2004109528A2 | Cites | World Intellectual Property Organization (WIPO) | Search report |
| US2004193821A1 | Cites | United States of America | Examiner |
| JPH04369720A | Cites | Japan | Examiner |
| US20040193821A1 | Cites | United States of America | – |
| JP04369720A | Cites | Japan | – |
| WO2004109528A2 | Cites | World Intellectual Property Organization (WIPO) | – |
| WO2004102403A2 | Cites | World Intellectual Property Organization (WIPO) | – |
13 members in 7 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 11047890 | United States of America | – | |
| 4789005 | United States of America | A | |
| 4789005 | United States of America | A | |
| 2006003445 | United States of America | W | |
| 2006003445 | United States of America | W | |
| 2005047890 | – | – | – |
| 2006003445 | – | – | – |
| US20050047890 | – | – | – |
| WO2006US03445 | – | – | – |
Members13
| Document | Office | Kind | |
|---|---|---|---|
| WO2006083899A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2006195631A1 | United States of America | A1 | |
| TW200641621A | Taiwan Province of China | A | |
| GB0714910D0 | United Kingdom | D0 | |
| KR20070092318A | Republic of Korea | A | |
| GB2438116A | United Kingdom | A | |
| GB2438116A8 | United Kingdom | A8 | |
| DE112006000298T5 | Germany | T5 | |
| JP2008529175A | Japan | A | |
| US2009013108A1 | United States of America | A1 | |
| TWI335514B | Taiwan Province of China | B | |
| JP4891925B2This record | Japan | B2 | |
| US8166218B2 | United States of America | B2 |
19 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 4891925
- Publication, DOCDB
- 4891925
- Publication, EPODOC
- JP4891925B
- Application
- 2007553363
- Application, DOCDB
- 2007553363
- Application, EPODOC
- JP20070553363
Titles2
- Japanese
- メモリモジュールからローカルデータをマージするためのメモリバッファ
- English
- Memory buffer for merging local data from memory modules
Classification
- CPC, 15
- G11C7/1078
- G06F5/06
- G06F12/00
- G06F13/1684
- G11C5/04
- G11C7/10
- G11C7/1051
- G11C7/222
- G11C11/4093
- G11C2207/107
- G06F7/74
- G06F13/1647
- G06F13/1673
- G11C7/103
- G11C7/1036
- IPC, 3
- G06F13 16
- G06F12 04
- G06F12 00