Method for concatenating frames in communication system
1 claim: 1 independent, 0 dependent
- 1ディジタル化されたオーディオ信号の伝送に関連して隠蔽サンプルのシーケンス を 生成する方法であって、 サンプルの時間順序でバッファされた、オーディオ信号のディジタル化された表現のサンプルの 複数の サブシーケンス か ら、隠蔽サンプルのシーケンス を 生成するステップを有し、 前記生成するステップは、 前記バッファされたサブシーケンス を 、最後にバッファされたサブシーケンスから始めて、時間の逆方向に 所定数のバッファされたサブシーケンスを 読み出した隠蔽サブシーケンスから構成されるステップバックシーケンス(逆順シーケンス)と、 前記バッファされたサブシーケンスを、 前記ステップバックシーケンス の最後のサブシーケンスと対応するバッファされたサブシーケンスに、前記バッファにおいて時間の順方向で続くバッファされたサブシーケンスから始めて 、時間の順方向に 最後にバッファされたサブシーケンスまで 読み出した隠蔽サブシーケンスから構成される読み取り長さシーケンス(時間順シーケンス)との組をつくる段階と、 前記ステップバックシーケンスと前記読み取り長さシーケンスとの組をつくる段階を 、隠蔽サンプルのシーケンスの長さに応じて 複数回くり返す段階とを有し、 前記くり返す段階において、ステップバックシーケンスをつくる際に所定数読み出されるバッファされたサブシーケンスの数を、くり返し回数の増加に応じて増加させることを 特徴とする方法。
93 paragraphs, as filed
The present invention relates to a telecommunications system. The present invention relates specifically to methods, devices and devices for compensating for signal packet loss and / or delay jitter and / or clock skew in order to improve signal transmission quality over wireless communication systems and packet-switched networks.
Modern telecommunications is based on the digital transmission of signals. For example, in FIG. 1, transmitter 200 collects audio signals from source 100. This source may be due to speech by at least one person and other sound sources collected by the microphone, or may be a voice signal storage system or generator system such as a text-to-speech synthesis or dialogue system. There is also. If the source signal is analog, it is converted to a digital representation using an analog / digital converter. The digital representation is subsequently encoded and placed in the packet according to a format suitable for digital channel 300. Packets are transmitted over digital channels. Digital channels typically have multiple layers of abstraction.
In the layer of abstraction in Figure 1, the digital channel takes a sequence of packets as input and sends the sequence of packets as output. Typically, due to channel degradation caused by noise, imperfections, and overloads within the channel, the sequence of packets output typically results in the loss of some packets, and the arrival of other packets. Contaminated by time delays and delay jitter. In addition, clock differences between transmitters and receivers can result in clock skew. The role of the receiver 400 is to decode the received data packet, convert the decoded digital representation from the packet stream and decode it into a digital signal representation, and further decode these representations into a signal sync (signal sync device). It is to convert to a decoded audio signal in a format suitable for output to 500. The signal sink may be at least one person whose decoded audio signal is presented, for example, by at least one speaker, or it may be an audio or audio storage system or an audio or audio dialogue system or recognition device. is there.
It is the responsibility of the receiver to accurately reproduce the signal that can be presented to the sink. If the sink directly or indirectly contains multiple human listeners, the purpose of the receiver is perceived by the person with respect to the acoustic signal from one source or multiple sources when presented to the human listener. The purpose is to acquire an audio signal representation that accurately reproduces the impression and information. To ensure this role of the receiver in the general case where loss, delay, and delay jitter degrade the packet sequence received by the channel, and the packet sequence is degraded due to the presence of clock skew. Efficient concealment is needed as part of the receiver subsystem.
As an example, Figure 2 shows one possible implementation of the receiver subsystem to fulfill this role. As this figure shows, incoming packets are stored in the jitter buffer 410, and the decoding and concealment unit 420 obtains the received and encoded signal representations from it and decodes these encoded signal representations. And by concealing, a signal representation suitable for storage in the reproduction output buffer 430 and subsequent reproduction output is obtained. Control of when to start hiding and what the specific parameters of hiding, such as the length of the signal to be hidden, is, may be performed by the control unit 440, for example. Here, the control unit 440 monitors the contents of the jitter buffer and the reproduction output buffer, and controls the operation of the decoding / concealment unit 420.
Concealment may also be achieved as part of the channel subsystem. FIG. 3 shows an example of a channel subsystem in which packets are forwarded from channel 310 to channel 330 via a subsystem 320, which is referred to later as a relay. In a real system, this relay function can be used for various types of routers, proxy servers, edge servers, network access controllers, wireless local area network controllers, voice over IP gateways, media gateways, unlicensed network controllers, unlicensed network controllers and others. It can be achieved by units called by various names that depend on the context, such as the name of. In the context of this specification, these are all examples of relay systems.
Figure 4 shows an example of a relay system that can hide audio. As shown in this figure, the packet is transferred from the input buffer 310 to the output buffer 360 via the packet switching subsystems 320 and 350. The control unit 370 monitors the input and output buffers and, as a result of this monitoring, decides whether transcoding and concealment are necessary. If necessary, the switch directs the packet through the transcoding and concealment unit 330. If not required, the switch directs the packet through the minimum protocol action subsystem 340. Here, the minimum protocol action subsystem 340 performs minimal actions on the packet header so that the packet follows the protocol to which it is applied. This may include changing the sequence number and time stamp of the packet.
When transmitting an audio signal using a system exemplified by, but not limited to, the above description, loss, delay, delay jitter and / or clock skew in the signal representing or partially representing the audio signal. Need to hide. Prior arts that approach this concealment task are classified into pitch repetition methods and time scale correction methods.
The pitch repetition method that may be embodied in the oscillator model is based on an estimate of the pitch period in the uttered speech, or an estimate of the corresponding fundamental frequency of the uttered speech signal. Given a pitch period, the concealed frame is acquired by repeating the reading of the final pitch period. Discontinuities between the beginning and end of the concealed frame and each iteration of the pitch period may be smoothed using a windowed overlap addition procedure. For example, refer to Patent Document 1 and Non-Patent Document 1 relating to the pitch repetition method. Prior art systems integrate pitch iteration-based concealment with a decoder based on linear predictive coding principles. In these systems, pitch repetition is typically achieved by long-term prediction or reading from an adaptive codebook loop in the linear prediction behavior domain. For concealment based on pitch repetition in the linear prediction motion domain, refer to, for example, Patent Document 2, Non-Patent Document 2 and Non-Patent Document 3. The methods described above apply to concealing lost or increasing delays, i.e. positive delay jitter, and underflow or near-underflow of inputs or jitter buffers due to, for example, clock skew. It is necessary to generate a shortened concealment signal in order to conceal a reduced delay, negative delay jitter or a situation near an input or jitter buffer overflow or overflow. The pitch-based method achieves this by an overlapping addition procedure between the pitch period and the earlier pitch period. Refer to Patent Document 1 as an example of this method.
<patcit num="1"><text>International Publication Patent No. 0148736 Pamphlet.</text></patcit><patcit num="2"><text>U.S. Pat. No. 5,699481.</text></patcit><nplcit num="1"><text>International Telecommunications Union recommendation ITU-T G.711 Appendix 1.</text></nplcit><nplcit num="2"><text>International Telecommunications Union recommendation ITU-T G.729.</text></nplcit><nplcit num="3"><text>International Engineering Task Force Request for Comments 3951.</text></nplcit><nplcit num="4"><text>Linag, Farber, Girod, Adaptive Playout Scheduling and Loss Concealing for Voice Communication over IP Networks, Multi IEEE Transactions on Multimedia, December 2003, Volume 5, Issue 4, p.532-p.543.</text></nplcit><nplcit num="5"><text>Rφdbro, Jensen, Time-scaling of Sinusoids for Intelligent Jitter Buffer in Packet Based Telephony, 2002, Voice Coding IEEE Proceeding Workshop on Speech Coding, p.71-p.73.</text></nplcit><nplcit num="6"><text>Valenzuela, Animalu, "A new voice-packet reconstruction technique", 1989, IEEE.</text></nplcit>
<p> This can also be achieved by leveraging the facilities that exist within the linear prediction decoder. As an example, Patent Document 2 discloses a method of simply discarding a particular codebook contribution vector from a reproduction signal, depending on the state of the adaptive codebook, in order to guarantee the periodicity of the pitch in the reproduction signal. .. One purpose related to the pitch repeating method is seamless signal continuity from the concealed frame to the next frame. Patent Document 1 discloses a method for achieving this object. According to the invention disclosed in Patent Document 1, this object is achieved by a concealed frame having a length that is time-denatured and possibly signal-dependent. While this solution can efficiently guarantee seamless signal continuity in relation to delay jitter and clock skew concealment, it has flaws with respect to the type of system depicted in FIG. That is, according to this type of concealment, it guarantees the coding of the concealment to a preset fixed length frame that seamlessly connects to the already encoded frame, which is preferably relayed via the minimum protocol action 340. Can not do it.</p><p> A frequent problem with pitch iteration-based methods for concealing losses and rapidly increasing delays is that the pitch cycle iterations make the reproduced signal voice unnatural. More specifically, this audio signal becomes too periodic. In the worst case, the so-called string sound (string) in the reproduced audio signal sounds) is perceived. Prior art has many ways to alleviate this problem. These methods include the use of repeating cycles that are twice or three times the estimated pitch period. As an example, Non-Patent Document 3 describes a method in which twice the estimated pitch period is used if the estimated pitch period is less than 10 milliseconds. As another example, Non-Patent Document 1 describes a method in which pitch cycle doubling and later doubling are introduced to repeat two pitch cycles and later three pitch cycles instead of repeating a single pitch cycle. Is described. See Non-Patent Document 1 for a complete description of this method. Further, in order to mitigate the string sound, a mixture of the concealed signal and a random or random signal component having a level depending on the vocalization level of the voice and a gradual attenuation of the concealed signal is introduced. Will be done. Sometimes this random signal is derived by arithmetic on the buffered signal or by using a facility such as a random codebook already available in the decoder. See Patent Document 2, Non-Patent Document 2 and Non-Patent Document 3 for examples of using such features. Gradual attenuation is also used to suppress the introduced artifacts. This may be the best option for the near-end listener to interpret, given the basic concealment method, but the far-end listener will return the echo and adapt this echo. The effect of this attenuation can be overwhelmingly negatively interpreted in the way the filter cancels. This is because the attenuation reduces the sustainability of the operation of the adaptive echo canceller. This reduces the tracking quality of this to the actual echo path, and far-end listeners may experience greater echo return.</p><p> For example, the type of timescale correction method described in Non-Patent Document 4 works through a matched smooth overlap addition procedure. In this procedure, the signal segment is buffered but not yet regenerated, the signal is bluntly windowed and identified as a template segment, followed by other bluntly windowed to identify similar segments. The segment is searched. Here, the similarity may be, for example, a correlation measure. Smoothly windowed template segments and smoothly windowed similar segments are subsequently overlapped and added to produce a timescale-corrected signal. When the playback time scale is extended, the search area for similar segments is positioned before the template segment in the sample time. Conversely, when the playback time scale is compressed, the search area for similar segments is positioned ahead of the template segment in sample time. In well-known timescale correction methods, template lengths and similar segments and the windows applied to them are predefined before the timescale correction is performed, and these amounts are for the particular signal to which this timescale correction is applied. Not adapted according to characteristics. As observed in Non-Patent Document 4, which uses time scale correction by the prior art, spike delay is start-time in low-delay playback scheduling as required for real-time two-way voice communication on a packet network. It cannot be effectively mitigated from position.</p><p> Other methods are known that have similarities to the time scale correction method and the pitch repetition method. One type to mention in this context is a sine wave-based concealment method. See, for example, Non-Patent Document 5. Depending on the amount of interpolation or pitch repetition achieved through the sinusoidal model domain by these methods, these methods have the same limitations identified with respect to the pitch repetition and timescale correction methods mentioned above. receive.</p>
<p> The disclosed inventions or embodiments thereof alleviate constraints previously identified in known solutions, such as audible artifacts, and other unspecified defects in the known solutions described above. ..</p><p> Specifically compared to a known pitch repetition based method, the disclosed method provides a technique for generating a concealed signal that represents an audio signal. Here, this concealed signal has significantly less perceptually annoying artifacts such as string sounds. As a result, the constraints of these systems are relaxed and the perceived speech quality is directly improved. This is also achieved at the same time as the introduction of significantly less attenuation in the concealed signal. This relieves the second constraint of the system based on repeated pitches. In addition, the relaxation of this second constraint directly improves the perceived quality of the hidden signal on the near end side of the communication. In addition, the relaxation of the second constraint improves the perceived quality at the far end of communication in systems with near-end acoustic echo and adaptive filters to mitigate the effects of acoustic echo perceived by the far end. .. This second effect is due to the disclosed concealment signals of the method, which exhibit less attenuation and provide more sustained operation for the adaptive process of the adaptive echo canceling filter. Achieved by Moreover, the robustness of the disclosed technology to acoustic background noise surpasses that of known pitch repeat-based methods.</p><p> In addition, the disclosed method, specifically compared to known timescale correction methods, has low latency playback or output buffer scheduling as required for real-time two-way voice communication over packet networks. Allows hiding of spike delays in the system. This relaxes the main constraints on known timescale modifications.</p><p> In a first aspect, the invention provides a method for generating a sequence of concealed samples in connection with the transmission of a digitized audio signal, a sample of the buffered audio signal of the digitized representation. The at least two consecutive subsequences of the sample within the sequence of the concealed sample are based on the subsequence of the buffered sample, comprising generating the sequence of the concealed sample from the sample in chronological order. The buffered sample subsequences are contiguous in the sorted time order.</p><p> The following definitions apply to the first aspect above and are used throughout this disclosure. The term "sample" originates from a sample originating from a digitized audio signal, or a sample originating from a signal derived from the digitized audio signal, or a coefficient or parameter representation of such a signal. These coefficients or parameters are scalar or vector values. The term "frame" is understood to be a set containing successive samples, using the above definition of samples. A "subsequence" is understood to be a set containing at least one contiguous sample, using the above definition of a sample. Therefore, in certain cases, the subsequence is equal to the sample. For example, in the case of using overlap addition, two consecutive subsequences may contain multiple overlapping samples. Depending on the frame selection, the subsequence may extend between two consecutive frames. In a preferred embodiment, the subsequences are arranged so that one subsequence cannot be a subset of another.</p><p> Preferably, at least two contiguous subsequences of the sample within the concealed sample sequence are based on the buffered sample subsequence, and the buffered sample subsequence is contiguous in reverse chronological order. Thus, in a preferred embodiment, the concealed sample sequence comprises a contiguous subsequence, such as a contiguous sample based on a contiguous buffer sample in the reverse chronological order. For example, two, three, four, or more contiguous subsequences of a sample in a concealed sample sequence may be based on a subsequence of a buffered sample that is contiguous in the reverse chronological order. In other words, the concealment sequence that occurs preferably includes a portion that is more or less based on direct reverse reproduction of the buffered sample. In one preferred embodiment, the concealed sample sequence comprises a continuous sample set of buffered samples in reverse chronological order. By calculating at least a portion of the concealed sample sequence based on the buffered sample using this sort or reverse sort method, it is more natural, unaffected by prior art string sounds. Pronouncing concealment sequences are provided, and the removal or reduction of some other artifacts is also facilitated.</p><p> The method described has many advantages in relation to communication systems, for example VoIP systems. Here, since the digital audio signal is transmitted in frames and the communication is exposed to frame loss and jitter, there is a need for a sample concealment sequence to at least partially reduce sudden changes in the audible and annoying signal.</p><p> In a preferred embodiment, the position of the buffered sample subsequence is located at points that gradually expand posteriorly and anteriorly in sample time during the generation of the concealed sample. This may be done by an index pattern generator that controls this temporal expansion. By analyzing the buffered sample, this index pattern generator selects the start, stop and velocity of the backward temporal expansion path, which also selects the start, stop and velocity of the forward temporal expansion. It also controls patterns that sequence backward and forward temporal expansions to generate a natural pronunciation concealment sequence.</p><p> The sequence of concealed samples may start with a subsequence based on the last subsequence in the time order of the buffered samples.</p><p> The time reordering of the subsequence may be based on a sequential process of reading and indexing the samples in the time direction and stepping in the opposite direction of the time. Preferably, the sequential process of indexing and reading the sample is a) A step of indexing buffered samples by stepping some buffered samples in the reverse order of time, followed by, b) Some buffered samples are read in chronological order, starting with the buffered samples indexed in step a), and the read samples are sub-sequences of the concealed samples. Including the steps used to calculate the sequence The number of samples read and buffered in the time direction is different from the number of buffered samples stepped in the opposite direction of the time. This difference in numbers avoids the periodicity that leads to unnatural string sounds. The method is further referred to as "backstep" and "read length" in a detailed description of later embodiments.</p><p> The number of buffer samples read in the time direction may be greater than or less than the number of buffer samples stepped in the opposite direction of time. Preferably, the number of samples read and buffered in the time direction is less than the number of buffer samples stepped in the opposite direction of the time. This selection provides a way to further expand in the buffered sample in the opposite direction of the gradual time, thus concealing sequences in which subsequent samples are based on buffer samples older than gradual and then forward expansion is initiated. I will provide a.</p><p> The subsequence of the hidden sample sequence may be calculated from the buffered sample subsequence by involving a weighted overlap addition procedure. The weighting function in the weighted overlap addition procedure may be a function of frequency. The weighted overlap addition procedure may be modified in response to the matching quality indicator. This matching quality indicator is a measure for two or more subsequences of the sample entered in the weighted overlap addition procedure above.</p><p> The time-wise sort may be partially described by the backward and forward expansion of the location pointer. Preferably, the backward expansion of the location pointer is limited by the use of stop criteria. The stop criteria for the posterior deployment, the pace (or speed) of the anterior and posterior deployments, and the number of the posterior deployments initiated so as to optimize voice quality when interpreted by a human listener. May be optimized at the same time.</p><p> Preferably, the smoothing and equalization operations are applied to the buffered sample. This may be done before the sample is buffered, during buffering, or just before the sample is used to calculate the hidden sample. The stop criteria for backward deployment, the pace of forward and backward deployment, the number of backward deployments initiated, and the smoothing and equalization operations determine the audio quality when interpreted by a human listener. It may be optimized at the same time to be optimized.</p><p> The backward and forward expansions of the location pointer may be simultaneously optimized to optimize voice quality when interpreted by a human listener.</p><p> Preferably, phase filtering is applied to minimize the discontinuity at the boundary between the sequence of concealed samples and the contiguous frames of the sample. The introduction of phase filtering facilitates the reduction of discontinuity, which is a well-known problem when introducing concealment sequences. When such phase filtering is applied, the simultaneous optimization also includes the signal distortion introduced by the phase filtering so as to optimize the speech quality as perceived by a human listener. May be good.</p><p> Noise mixing may be introduced into the sequence of concealed samples. In particular, noise mixing may be introduced into the sequence of concealed samples, which is modified in response to a sequential process of reading and indexing the sample in the time direction and stepping in the opposite direction of time. .. In such cases, the sequential process of reading and indexing the sample in the time direction and stepping in the opposite direction of time and the response to it may include the use of matching quality indications.</p><p> An decay function may be applied to the sequence of concealed samples. In particular, such decay functions may be modified in response to a sequential process of reading and indexing samples in the time direction and stepping in the opposite direction of time. A sequential process of reading and indexing samples in the time direction and stepping in the opposite direction of time and the response to this may include the use of matching quality indications.</p><p> Preferably, the final number of samples in the concealed sample sequence is preset, for example, the number of samples in the concealed frame may be fixed. The number of samples is preferably independent of the characteristics of the digital audio signal. The preset number of samples is a preset integer value in the range of 5-1000, such as in the range of 20-500, and preferably depends on the actual sample frequency.</p><p> The sequence of the concealment sample may be included in the first concealment frame. The method may further include generating at least one second concealment frame following the first concealment frame, the second frame comprising a sequence of second concealment samples. The sequences of the concealed samples in the first and second concealed frames are preferably different, i.e., consecutive copies of both concealed frames are preferably avoided. The use of frames containing different concealment sequences leads to more natural vocal concealment. Preferably, the first and second concealment frames contain the same number of samples.</p><p> Preferably, at least one subsequence of the sample in the second concealment frame becomes a subsequence of the sample further buffered in the opposite direction of time than any subsequence of the sample contained in the first concealment frame. At least partially based. Therefore, the hiding frame that comes after is preferably based on an older buffer sample.</p><p> In a second aspect, the invention provides program code that can be executed by a computer adapted to perform the method according to the first aspect. Such program code may be written in a machine-dependent or machine-independent format and in any programming language, such as machine code or a higher programming language.</p><p> In a third aspect, the present invention provides a program storage device comprising an instruction sequence for a microprocessor such as a general purpose microprocessor for performing the method according to the first aspect. The storage device may be any type of data storage means such as a disk, a memory card or a memory stick, a hard disk, or the like.</p><p> In a fourth aspect, the invention provides a device, eg, a device or device, for receiving a digitized audio signal. -A memory means for storing a sample representing a received digital audio signal, -Includes processor means for performing the method according to the first aspect above.</p>
<p> To implement the present invention with appropriate means, such as those described below in connection with a preferred embodiment, the decoder and concealment system and / or the transcoder and concealment system introduce perceptually annoying artifacts. It makes it possible to efficiently hide the sequence of packets that are lost or delayed. Furthermore, this is achieved without introducing high speed fading, with acoustic background noise and robustness to multiple speakers. The improvement in robustness is achieved by the consistency of the method being less dependent on strict signal periodicity over time evolution than on iteration-based methods. Thereby, the present invention enables high quality two-way voice communication in situations with acoustic background noise, acoustic echo and / or severe clock skew, channel loss and / or delay jitter.</p>
Next, the present invention will be described in more detail with reference to the accompanying drawings.
The present invention can take various modifications and alternative forms, but the drawings show specific embodiments by way of example. Hereinafter, these specific embodiments will be described in detail. However, it should be understood that the present invention should not be limited to these particular forms disclosed. Rather, the invention includes all modifications, equivalents and alternatives within the spirit and scope of the invention as defined by the appended claims.
The method according to the invention is invoked in the decoding and concealment unit 420 of the receiver, such as that shown in FIG. 2, or in the transcoding and concealment unit 330, such as that shown in FIG. 4, or its actions are appropriate. Launched at any other location in the communication system. At these locations, some frames of the buffered signal are available and some concealment frames are required. The available signal frame and the required concealment frame may consist of, for example, a time domain sample of an audio signal, which is an audio signal, or a sample such as a linear predictive motion sample derived from the above sample. It may consist of other coefficients derived from the audio signal that represent the audio signal frame completely or partially. Examples of such coefficients include frequency domain coefficients, sinusoidal model coefficients, linear predictive coding coefficients, waveform interpolation coefficients and other set of coefficients that represent the audio signal sample completely or partially.
FIG. 5 shows a preferred embodiment of the present invention. According to FIG. 5, the available signal frame 595 is stored in the frame buffer 600. The signal frame 595 is a frame received and decoded or transcoded, or a concealed frame from an earlier operation by this method or other method for generating a concealed frame, or a combination of the above types of signal frames. It may be. The signal in the framebuffer is analyzed by the index pattern generator 660. The index pattern generator can effectively utilize the estimated values of signal pitch 596 and utterance 597. Depending on the overall system design, these estimates may be available as inputs from other processes such as coding, decoding or transcoding processes, or by other methods, preferably signals. Calculated using state-of-the-art methods for analysis. In addition, the index pattern generator employs as inputs 598 hidden signal frames to generate and a pointer 599 pointing to the beginning and end of at least one particular signal frame replaced by the hidden frames in the framebuffer. As an example, if these buffers point to the end of the framebuffer, this means that at least one hidden frame should be made to follow the signal stored in the framebuffer. As another example, if these pointers point to a non-empty subset of consecutive frames in the framebuffer, then this means that at least one hidden frame represents or partially represents the audio signal in the frame sequence. It means that it should be made to replace the frame it represents.
To further illustrate this, it is assumed that the frame buffer 600 includes signal frames A, B, C, D, E and the number of concealed frames 598 is 2. Then, if the pointer 599 pointing to the frame to be replaced points to the end of the framebuffer, this means that the two hidden signal frames should be made to follow the signal frame E in turn. Conversely, if the pointer 599 points to signal frames B, C, and D, then these two concealed frames will replace signal frames B, C, and D, and in turn follow signal frame A, and in turn. It should be made so that the signal frame E follows.
With respect to the number of concealed frames 598 and the method of determining the subset of frames that the concealed frames should ultimately replace, i.e. the pointer 599, the state-of-the-art method should preferably be used. Thus, data 596, 597, 598 and 599 and signal frame 595 constitute inputs to the methods, devices and devices according to the invention.
In a given overall system design, the length or size of the signal frame is effectively maintained as a constant during the execution of the concealment unit. This is a typical case among other methods when the concealment unit is integrated into the relay system. Here, in the relay system, the result of concealment should be contained in a packet representing an audio signal within a preset length of time, and this preset length is determined elsewhere. Will be done. As an example, this preset length may be determined during protocol negotiation during call setup in a voice-over IP system, and may be modified during the conversation, eg, in response to a network congestion control mechanism. May be good. As will become apparent later, some embodiments of the present invention meet this requirement of effectively operating at a preset signal frame length. However, such innovations are not limited to these system requirements, and other embodiments of this innovation will also work with concealment of non-integer number of frames and concealment frames with time-varying lengths. These lengths can be a function of specific content in the framebuffer, perhaps in combination with other elements.
In the embodiment of the present invention, the smoothing and equalization operation 610 acting on the signal 605 from the frame buffer can be effectively utilized. This smoothing and equalization is a signal with increased similarity to at least one signal frame in which a frame temporally earlier than at least one concealment frame is replaced by the at least one concealment frame or a frame immediately preceding it. Generate 615. Alternatively, if at least one concealment frame is inserted into a sequence having an existing frame without substitution, the similarity is with the similarity to at least one frame immediately preceding the intended position of the at least one concealment frame. Become. For later reference, both of these cases are simply referred to as similarities. Similarity is the similarity when interpreted by a human listener. Smoothing and equalization obtains a signal with increased similarity, but at the same time maintains the natural vocal development of signal 615. Examples of similarity-increasing operations effectively performed by Smoothing and Equalization 610 are parameters such as energy envelopes, pitch contours, speech grades, speech cutoffs, spectral envelopes and other perceptually important parameters. Includes increased smoothness and similarity in.
For each of these parameters, the abrupt transitions of the parameter expansion in the frames to be smoothed and equalized are filtered out, and the average parameter level in these frames is in a similar sense as defined above. It is smoothly modified to be more similar. Effectively, similarity is introduced only to the extent that the signal development of natural vocalizations is still maintained. Under the control of the index pattern generator 660, smoothing and equalization can effectively mitigate transitions and discontinuities that would otherwise occur in the next indexing and interpolation operation 620. In addition, pitch contour smoothing and equalization is effective in minimizing distortion that is otherwise introduced into the concealed frame by the index pattern generator 660 and ultimately by the phase filter 650 later. It may be controlled. Smoothing and equalization operations are effective in replacing, mixing, interpolating and / or merging signals or parameters with signal frames (or their parameters derived) further found in the reverse direction of time in the framebuffer 600. Can be used for. The smoothing and equalization operation 610 may be excluded from the system without departing from the general scope of the invention. In this case, the signal 615 will be equated with the signal 605, and the signal input 656 and control output 665 of the index pattern generator 660 may be omitted from the system design.
The indexing and interpolation operation 620 takes as input the probably smoothed and equalized signal 615 and index pattern 666. Moreover, in some effective embodiments of the present invention, the indexing and interpolation operations take the matching quality indicator 667 as input. The matching quality indicator may be a scalar value per hour or a function of both time and frequency. The purpose of the matching quality indicator will become apparent later in the text of this specification. The index pattern 666 parameterizes the operations of the indexing and interpolation functions.
FIG. 5A shows an example of how the index pattern can index subsequences in buffered samples BS1, BS2, BS3, BS4 in the opposite direction of gradual time in the synthesis of at least one concealed frame. In the illustrated example, contiguous subsequences CS1, CS2, CS3, CS within concealed frames CF1, CF2, CF3<u style="single">4</u>, CS5, CS6, CS7 are based on the buffered subsequences BS1, BS2, BS3 and BS4 of the samples in frames BF1, BF2. As can be seen from the figure, the concealment subsequence CS1-CS7 is represented by the functional notation CS1 (BS4), CS2 (BS3), CS3 (BS2), which means that CS1 is based on BS4, etc. Indexed from the subsequences BS1-BS4 buffered with the location pointer in the reverse direction and then in the gradual time direction. Thus, FIG. 5A serves as an example showing how contiguous subsequences within a concealed frame can be based on a contiguous buffered subsequence, but reordered in time and continued to each other. As can be seen from the figure, the first four concealed subsequences CS1 (BS4), CS2 (BS3), CS3 (BS2) and CS4 (BS1) are the four subsequences BS1, BS2, BS3 at the end of the buffered sample. , BS4 are selected in a contiguous order, but in the reverse chronological order, so that they are based on the last buffered subsequence BS1 as the starting point. The first four subsequences of the reverse time sequence are followed by three consecutive buffered subsequences, all based on BS2, BS3 and BS4, respectively, in time sequence CS5, CS6, CS7. This preferred index pattern is the result of the index pattern generator 660 and can vary significantly with inputs to this block 656, 596, 597, 598 and 599. FIG. 5B represents another example exemplifying how the concealed subsequence CS1-CS11 can be generated based on the temporal rearrangement of the subsequences BS1-BS4 buffered according to the notation in FIG. 5A. As can be seen, the time-slow concealment subsequences are based on progressively, further buffered subsequences in the opposite direction of time. For example, the first two consecutive concealment subsequences CS1 and CS2 are the most The latter two buffered subsequences BS3 and BS4 are based on the reverse time order, while the time-slow concealment subsequences such as CS10 are used to calculate BS1, ie CS1 and CS2. Based on more buffered subsequences in the opposite direction of time. Thus, Figure 5B shows that successive concealment subsequences are based on buffered subsequences that are indexed back and forth in time in such a way that indexing expands in the opposite direction of gradual time. is there.
In an effective embodiment of the invention, this gradual evolution of time in the reverse direction is of a sequence of stepbacks referred to as intent herein, and a read length referred to as intent herein. Formalized as a sequence. In a simple embodiment of the index pattern of this format, the signal sample or the pointer to the parameter or coefficient representing the signal sample is moved backward by an amount equal to the first step back, after which a certain amount is placed in the concealed frame. A sample or a parameter or coefficient representing the sample is inserted. The above quantity is equal to the first read length. After this, the pointer is retracted by an amount equal to the second stepback, a sample amount equal to the second read length, or a parameter or coefficient representing the sample amount is read, and so on.
Figure 5C shows an example of this process in which the first count data of the indexed sample is sorted. This first count data is written on the signal time axis, while the count data written on the concealment time axis of FIG. 5C is sorted according to the placement of the original sample in its concealment frame. Correspond. In the case of this illustrated example, the first, second and third stepbacks are arbitrarily selected as 5, 6 and 5, respectively, and the first, second and third read lengths are similarly selected, respectively. Arbitrarily selected as 3, 4, or 3. In this example, the subsequences with the time index sets {6,7,8}, {3,4,5,6} and {2,3,4} are each subsequences that gradually expand in the opposite direction of time. is there. In this case, the stepback and read length sequences are chosen purely for illustration purposes. For example, in the case of audio residual samples sampled at 16kHz, the typical stepback value is in the range of 40 to 240, but is not limited to this range, and the typical reading length is in the range of 5 to 1000 samples. However, it is not limited to this range. In a more advanced embodiment of this format, from a forward-looking sequence (eg, a subsequence indexed in the original time direction or in the reverse direction of time) to another forward-looking sequence that takes one more step in the reverse direction of time. The transition is made incrementally by a gradual shifting interpolation.
FIG. 6 shows the operation of a simple embodiment of the indexing and interpolation functions in response to one step back and the corresponding read length and matching quality indicators. Here, for the purpose of mere illustration, the signal frame consists of time domain audio samples. Gradually shifting interpolation is based on the general definition of the term "sample" as used herein, i.e., including coefficients or parameters of scalar or vector values that represent a time domain audio sample. Similarly, it is therefore applied directly. In this figure, 700 indicates a segment of signal 615. Pointer 705 is the sample time following the sample time of the last generated sample in the indexing and interpolation output signal 625. The time interval 750 has a length equal to the read length. The time interval 770 also has a length equal to the read length. The time interval 760 has a length equal to the step back. The signal sample starting at time 705 at 700 and the temporally forward read length are multiplied one by one by the window function 720. Similarly, the signal sample starting from the location before the location 706 after the step back of one sample in 700 and the sample of the read length after that are also multiplied one by one by the window function 710. The resulting samples from multiplication with window 710 and multiplication with window 720 are added one by one to 730, resulting in 740 forming a new sample batch of output 625 from indexing and interpolation operations. .. Upon completion of this operation, pointer 705 moves to location 706.
In a simple embodiment of the invention, the window functions 710 and 720 are simple functions with a read length of 750. One such simple function selects window 710 and window 720 as the first and second halves of the Hanning window, which are twice the read length, respectively. In this case, a wide range of functions can be selected, but in view that such functions must be meaningful in the context of the present invention, these are the samples in the segment shown by 750 and 770. Weighted interpolation must be achieved by gradually, but not necessarily monotonically, moving from the high weight for the segment indicated by 750 to the high weight for the segment indicated by 770 to and from the sample shown.
In another embodiment of the invention, the window functions 710 and 720 are functions of the matching quality indicator. In a simple example of such a function, the interpolation operation sums in either amplitude or power, depending on the normalized correlation threshold on the segment of the signal 700 represented by the time intervals 750 and 770. Is selected to be 1. Another example of such a function optimizes the window weight as a matching measure-only function, instead of avoiding the constraint of summing the amplitude or power to 1. A further improvement on this method is to find the actual value of the normalized correlation and in response, optimize the interpolation operation using, for example, a classical linear estimation method. Examples of preferred methods will be described later, but in these examples the normalized correlation threshold or actual value is an example of effective information sent by the matching quality indicator 667. According to a preferred embodiment described below, the interpolation operation may have different weights implemented at different frequencies. In this case, the matching quality indicator 667 can effectively send the matching measure as a function of frequency. In an effective embodiment, this weight as a function of frequency is implemented as a multi-stage delay line or as another parametric filter form that can be optimized to maximize the matching criteria.
FIG. 6 shows indexing and interpolation operations when the signal 615 (and thus the signal segment 700) contains a sample representing a time domain sample of the audio signal or of the time domain signal derived from the audio signal. It is shown. As mentioned above, the samples at frame 595 and thus at signals 605 and 615 may effectively be such that each sample is a vector (vector value sample). Such vectors include coefficients or parameters that represent or partially represent the audio signal. Examples of such coefficients are the frequencies of the line spectrum, the frequency domain coefficients, or the coefficients that define a sinusoidal signal model such as a set of amplitudes, frequencies and phases. Based on a detailed description of preferred embodiments of the present invention, the design of interpolation operations effectively applied to vector-valued samples can be found by reading the general literature on individual specific cases of such vector-valued samples. Since the details of the above are also described, it is feasible for those skilled in the art.
In understanding the present invention, if the indexing and interpolation operations are repeatedly performed with a read length smaller than the stepback, the sample at signal 625 will eventually be advanced gradually and in the opposite direction at signal 615. It is effective to notice that it is representative of signal samples. Thus, if the stepback and / or read length is changed so that the read length is longer than the stepback, this process is reversed so that the sample at signal 625 gradually advances at signal 615. It is a representative of signal samples that can be advanced in the time direction. With effective selection of step-back sequences and read-length sequences, long concealed signals with abundant and natural variants can be sampled in time preceding from the last received signal frame in framebuffer 600. Obtains a sample that precedes another preset time, either without needing or that can be positioned earlier than the last sample in the last received frame in framebuffer 600. be able to. As a result, the present invention makes it possible to conceal delay spikes in systems with low delay playback or output buffer scheduling. In the formulation of this specification, the simple and exact reverse time evolution of the signal, which may be useful to consider as an element in a simple embodiment of the invention, is a reading of one sample. This is achieved by repeated use of length, step back of two samples, and window 720 consisting of a single sample with a value of 0 and window 710 consisting of a single sample with a value of 1.0.
The main purpose of the index pattern generator 660 is to control the actions of the indexing and interpolation operation 620. In a series of preferred embodiments, this control is formalized into an indexing pattern 666, which may consist of a sequence of stepbacks and a sequence of read lengths. This control may be further extended with a sequence of matching quality indications, each of which may be a function of frequency, for example. An additional feature that may be output from the index pattern generator and whose use will become apparent later herein is 668 iterations. The number of iterations means the number of times at least one concealment frame assembly begins to unfold in the opposite direction of time. The index pattern generator converts these sequences into a smoothing and equalization signal 656 output from the smoothing and equalization operation 610, pitch estimation 596, vocalization estimation 597, number of concealed frames to be generated 598, and frames to be replaced. Acquires based on information that may include a pointer to 599 pointing to. In one embodiment of the index pattern generator, the generator enters different modes depending on the vocal indicator. Hereinafter, such a mode will be illustrated.
As an example effectively used in the linear prediction motion domain, the vocal indicator robustly indicates that the signal is unvoiced or that there is no active speech in the signal, i.e. the signal consists of background noise. For example, the index pattern generator can enter a mode in which a simple reversal of the temporal expansion of the signal sample is initiated. As mentioned above, this may be achieved, for example, by submitting a sequence with a stepback value of 2 and a sequence with a read length value of 1. Based on the design option of itself identifying these values and applying the appropriate window functions as described above). In some cases, this sequence may continue until the reverse temporal expansion of the signal is implemented for at least half the number of new samples required for one concealment frame, after which the value in the stepback sequence is 0. This may initiate a temporal expansion of the signal forward and continue until the pointer 706 effectively returns to the starting point of the pointer 705 in the first stepback application. However, this simple procedure is not always sufficient for high quality concealed frames. An important role of the index pattern generator is to monitor proper outage criteria. In the above example, the reverse temporal expansion may return the pointer 706 to a position in the signal where the audio, as interpreted by a human listener, is significantly different from the starting point. The temporal evolution should be reversed before this happens.
Preferred embodiments of the present invention can be applied to a set of stop criteria based on a series of measures. Some of these measures and stop criteria are illustrated below. If the vocalization indicates that the signal at the pointer 706 is voiced, then in the above example starting from unvoiced, the temporal expansion direction may be effectively reversed, as well as the pointer 706. The temporal expansion direction may be effectively reversed if the signal energy in the region surrounding is different from the signal energy at the starting point of pointer 705 (according to determination by absolute or relative threshold). .. As a third example, the spectral difference between the region around the starting point of pointer 705 and the current position of pointer 706 may exceed the threshold and the temporal expansion direction should be reversed.
A second mode example can be aroused when the signal cannot be robustly determined to be unvoiced or do not contain active voice. In this mode, the pitch estimate 596 is the basis for determining the index pattern. One step to do this is maximally normalized between the signal one pitch cycle ahead of the pointer 705 in time and the signal one pitch cycle ahead from the point earlier than the pointer 705 on stepback. Each step back is searched to give a good correlation. The search for the stepback value may be effectively limited to a certain area. This region may effectively be set to plus or minus 10 percent of the previously discovered stepback, or to the pitch lag if no such stepback has been discovered. Once the stepback is determined, the read length value determines whether the temporal signal expansion should be expanded in the opposite direction of time or in the time direction, and the execution speed of this expansion. Slow deployment is achieved by choosing a read length that is close to the stepback identification value. High speed deployment is achieved by choosing a read length that is much smaller or much larger than the step back for backward and forward deployments, respectively. The purpose of the index pattern generator is to select the read length to optimize the speech quality interpreted by the human listener. Choosing a read length that is too close to stepback can result in perceptually annoying artifacts such as string sounds, depending on the signal, such as a signal that is not sufficiently periodic. Choosing a read length that is too far from the stepback means that the larger time interval in the framebuffer is eventually swept during the temporal expansion of at least one hidden frame, or of temporal expansion. It implies that the orientation must be reversed more often until a sufficient amount of sample is generated for at least one concealment frame.
In the first case, some signals, such as signals that are not sufficiently stationary (or sufficiently smooth and not equalized), will eventually have some degree of similarity to stuttering in the audio of at least one concealed frame. May produce some kind of perceptually annoying artifacts. In the second case, artifacts such as string sounds may occur. One feature of an effective embodiment of the invention is that the read length can be determined as a function of stepback and normalized correlation. Here, the above function is optimized in the search for the optimum step back. One simple but effective option of this function in embodiments of the present invention is an example when this function acts on an audio signal and the signal frame contains a 20 ms linear predictive motion signal sampled at 16 kHz. Is given by the following function.
[Number 1] ReadLength = [(0.2 + NormalizedCorrelation / 3) * StepBack]
Here, square brackets [] are used to refer to rounding to the nearest integer, and the symbols ReadLength, NormalizedCorrelateion, and StepBack are used to refer to the read length and normalized correlation obtained for optimal stepback, respectively. And used to represent the corresponding stepback. The functions described above are included as merely examples to convey one effective option in some embodiments of the invention. The reading length options include any functional relationship that achieves this reading length, all of which are possible without departing from the spirit of the present invention. Specifically, an effective way to select the read length is to smooth and smooth using control 665 so that artifacts such as stuttering and string sounds reach their minimums at the same time in the intermediate concealment frame 625. Includes parameterizing the equalization operation 610. This explains why the index pattern generator 660 adopts the intermediate signal 656 instead of the output 615 as an input for smoothing and equalization operations, where the signal 656 is the final signal 615 controlled by the control 665. It represents a potential version and allows the index pattern generator to engage in optimization tasks through iterations. As with the previous silent and inactive voice modes, stop criteria are essential in this mode as well. All of the stop criteria proposed in the previous mode also apply to this mode. Moreover, in this mode, the stop criteria from the measurements for pitch and normalized correlation may effectively be part of an embodiment of the invention.
Figure 7 illustrates an effective decision theory for combining stop criteria. The quotation marks in FIG. 7 are as follows.
800: Identifies whether the signal is of high correlation type, low correlation type, or neither. Determine the initial energy level. 801: Determine the next step back and normalized correlation, and read length. 802: Determines if the signal has entered the low correlation type. 803: Determine if the signal has entered the highly correlated type. 804: Is the signal highly correlated type? 805: Is the signal a low correlation type? 806: Is the energy less than the relative minimum threshold or above the relative maximum threshold? 807: Is the normalized correlation below the high correlation type threshold? 808: Is the normalized correlation above the low correlation type threshold? 809: Did enough samples be generated?
For operations in the linear prediction motion domain of speech sampled at 16 kHz, the thresholds listed in FIG. 7 may be effectively chosen as follows. That is, the high correlation type may be entered when a normalized correlation greater than 0.8 occurs, and the threshold for staying in the high correlation type may be set to 0.5 for the normalized correlation. Often, the low correlation type may be entered when a normalized correlation less than 0.5 is spoken, and the threshold for staying in the low correlation type may be set to 0.8 with the normalized correlation. Often, the minimum relative energy may be set to 0.3 and the maximum relative energy may be set to 3.0. Moreover, in the context of the present invention, other logic and other stopping criteria may be used without departing from the spirit and scope of the present invention.
The application of the stop criteria is done in the opposite direction of time and then again in the time direction until sufficient samples are generated, or until the stop criteria are met. In a single deployment, the number of samples required for the concealment frame. It means that it is not guaranteed to bring. Therefore, another expansion performed in the opposite direction of time and in the time direction may be applied by the index pattern generator. However, if there are too many back-and-forth developments, some signals may produce artifacts such as string sounds. Therefore, preferred embodiments of the present invention are a stop criterion, a function applied to the calculation of the read length, a smoothing and equalization control 665 and the number of expansions back and forth, i.e. the number of iterations 668, and a pointer 599 pointing to the replacement frame. Further, if enabled by, the number of samples to be deployed in the time direction can be optimized simultaneously before each new deployment in the reverse direction of time is started. For this purpose, the smoothing and equalization operations may also be effectively controlled to slightly modify the pitch contour of the signal. In addition, this simultaneous optimization can take into account the operation of the phase filter 650 and slightly pitch contours to provide an index pattern that minimizes the distortion introduced into the phase filter at the same time as the other parameters described above. Can be changed. Based on the description of preferred embodiments of the present invention, one of ordinary skill in the art can understand that a variety of common optimization tools apply to this task. These tools include iterative optimization, Markov decision process, Viterbi algorithm, and so on. Any of these can be applied to this task without departing from the scope of the present invention.
Figure 8 is a flow graph showing an example of an iterative procedure that achieves simple yet efficient optimization of these parameters. The quotation marks in FIG. 8 are shown below.
820: Start controlling smoothing and equalization 665. 821: Acquire a new smoothing signal 656. 822: Activate the stop criterion. 823: Invokes the allowed number of iterations. 824: The index pattern of the back-and-forth expansion sequence that is evenly distributed over the available frames pointed to by pointer 599, or the time immediately following expansion in the time direction if the end of the available frame is specified. Identifies the index pattern of the sequence of expansions in the opposite direction of. 825: Is enough sample generated for the number of hidden frames 598? 826: Has the maximum number of iterations been reached? 827: Increase the number of repeat permissions. 828: Have you reached the loosest threshold of stop criteria? 829: Loosen the stop criteria threshold. 830: Change controls to increase the effects of smoothing and equalization.
If sufficient signals were not synthesized in at least one preceding temporal anterior expansion, then one temporal anterior expansion followed by one temporal anterior expansion may be effectively different. Please note. As an example, the sequence of stepbacks, read lengths and interpolation functions and the end location pointers after temporal anteroposterior expansion should be devised to minimize periodic artifacts that would otherwise result from repeating similar index patterns. Is. Taking the residual region sample of speech uttered at 16kHz as an example, one temporal anteroposterior expansion that produces, for example, about 320 samples is preferably about 100 more temporal anteroposterior expansions in the signal than the early temporal anteroposterior expansion The minute sample may be terminated retroactively in the opposite direction of time.
The embodiments disclosed so far effectively alleviate the problems of artificially generated string sounds known from the prior art method, while simultaneously reducing abrupt delay jitter spikes and abrupt repetitive packet loss. Allows for efficient concealment. However, in adverse network conditions, such as those encountered in some wireless systems, wireless ad hoc networks, best effort networks and other transmission methods, even the disclosed method may in some cases be within a concealed frame. In some cases, a slight component of tone control may be introduced. Therefore, in some embodiments of the present invention, the trace noise mixing operation 630 and the graceful attenuation filter 640 may be effectively applied. General techniques for mixing and attenuating noise are well known to those of skill in the art. This includes the effective use of the frequency-dependent time expansion of the power of the noise component and the frequency-dependent time expansion of the decay function. A unique feature of the use of noise mixing and damping in the context of the present invention is the manifestation of indexing patterns 666, matching quality measures 667 and / or iterations 668 for adaptively parameterizing noise mixing and damping operations. It is in use. Specifically, the index pattern indicates where the invariant signal sample is placed in the concealed frame and where the concealed frame sample is the result of the interpolation operation. In addition, the ratio of step back to read length, in combination with the matching quality measure, indicates the perceptual quality that results from the interpolation operation. Therefore, effectively, there is little or no noise that can be mixed with the original sample. Further noise may be effectively mixed with the samples that are the result of the interpolation process, and effectively the amount of noise mixed with these samples is effectively frequency-discriminatory matching. It may be a function of quality measure. In addition, the read length value for stepback also indicates the amount of periodic that can occur, and noise mixing may effectively include this measure in determining the amount of noise to mix with the concealed signal. This same The principle also applies to attenuation, which effectively uses graceful attenuation, but less attenuation may be introduced in the sample representing the original signal, which in the sample resulting from the interpolation operation. The above attenuation may be introduced. Further, effectively, the attenuation in these samples may be effectively a function of frequency-discriminatory matching quality indication. Again, the read length value for stepback indicates the amount of period that can occur, and the attenuation calculation may effectively include this measure in the design of the attenuation.
As mentioned in the background description of the invention, an important object of the subset of embodiments of the invention is to achieve a concealed frame of a preset length equal to the length of a normal signal frame. If this is desired from a system perspective, the means for this may effectively be a phase filter 650. Computationally simple and approximate operations for this block, but often sufficient, are smooth overlap addition between samples that exceed a preset frame length and tracking of samples from frames following a hidden frame. Achieving multiplication with the number of hidden frames that have a subset. Seen alone, this method is well known from the latest technology and is used, for example, in Non-Patent Document 1. In practice from a system perspective, this simple overlap addition procedure may be improved by multiplying the number of subsequent frames by -1 whenever it increases the correlation in the overlap addition region. However, for example, in transitions between voiced signal frames, other methods may be effectively used to further mitigate the effects of discontinuities at the frame boundaries. One such method is resampling of hidden frames. Seen as an independent method, this is also well known from the latest technology. See, for example, Non-Patent Document 6. Therefore, one of ordinary skill in the art can perform discontinuity relaxation at the frame boundary. However, in a preferred embodiment of the invention disclosed herein, resampling can effectively continue to the frame following the last concealed frame. This makes it possible to make the temporal changes that result from resampling techniques, and thus the gradient of the frequency shift, imperceptible to human listeners. Furthermore, the present invention is not resampling, but total time denaturation.<u style="single">Area</u>Disclose that a time-varying all-pass filter is used to mitigate discontinuities at frame boundaries. One embodiment is given by the following filter equation.
[Number 2] H_L (z, t) = (alpha_1 (t) + alpha_2 (t) * z ^ (-L)) / (alpha_2 (t) + alpha_1 (t) * z ^ (-L))
The function will be described below. Even if the sweep from the delay of L samples to the delay of 0 samples includes all or part of the samples in all or part of the concealed frame in the frame before the concealed frame and the frame after the concealed frame. Assuming desired over a good sweep interval, at the beginning of the sweep interval alpha_1 (t) is set to zero and alpha_2 (t) is set to 1.0 to provide a delay of L samples. .. As the sweep on t begins, alpha_1 (t) gradually increases to 0.5 and alpha_2 (t) gradually decreases to 0.5. When alpha_1 (t) equals alpha_2 (t) at the end of the sweep interval, the filter H_L (z, t) introduces zero delay. Conversely, the sweep from the delay of 0 samples to the delay of L samples is all or part of the samples in all or part of the concealed frame in the frame before the concealed frame and in the frame after the concealed frame. At the beginning of the sweep interval, alpha_1 (t) is set to 0.5 and alpha_2 (t) is set to 0.5 to provide a delay of 0 samples, if desired over the sweep interval. To. As the sweep on t is started, alpha_1 (t) gradually decreases to 0, and alpha_2 (t) gradually increases to 1.0. At the end of the sweep interval, when alpha_1 (t) reaches 0 and alpha_2 (t) reaches 1.0, the filter H_L (z, t) introduces a delay of L samples.
The filtering described above is simple to calculate but has a non-linear phase response. For perceptual reasons, this non-linear phase limits its use to the relatively small L. Effectively, L <10 for audio with a sampling rate of 16 kHz. One way to achieve filtering for larger initial values L is to activate several filters for multiple smaller values L that add up to the desired value L. Some of these filters can effectively be activated at different moments and sweep over different time intervals in their alpha region. Next, we disclose another way to increase the applicable range of L for this filter. A structure that provides the same filtering function as the method described above has L signals.<u style="single">Polyphase</u>Divide into these<u style="single">Polyphase</u>Perform the following filtering in each of.
[Number 3] H_1 (z, t) = (alpha_1 (t) + alpha_2 (t) * z ^ (-1)) / (alpha_2 (t) + alpha_1 (t) * z ^ (-1))
In the case of the present invention<u style="single">Polyphase</u>Filtering is effectively provided using upsampling. One way to do this effectively is with each<u style="single">Polyphase</u>Upsampled with a factor of K, and each upsampled<u style="single">Polyphase</u>In, filtering H_1 (z, t) is executed K times. Then, by downsampling with the coefficient K,<u style="single">Polyphase</u>The phase-corrected signal is reconstructed from. The coefficient K may be effectively selected as K = 2. The upsampling procedure obtains a near-linear phase response. This improves the perceptual quality interpreted by the human listener.
The above phase adjustments for multiple frames are applicable when hidden frames are inserted losslessly into the received frame sequence. This is also applicable when frames are extracted from the signal sequence to reduce the reproduction delay of subsequent frames. This can also be applied when frames are lost and zero or more concealed frames are inserted between frames received before and after this loss. In these cases, the method of acquiring the input signal of this filter and obtaining the delay L is as follows.
1) Continue or initiate the method disclosed herein or any other concealment method on a frame that is earlier in time than the discontinuity point. 2) Indexing time samples of L_test test samples into frames initiated by the method disclosed herein or any other concealment method, on frames that are later in time than discontinuity. Is reversed and inserted. 3) Apply a matching measure such as normalized correlation between at least one concealment frame from 1) and at least one frame from 2) containing L_test test samples that are headings. 4) Select L_test as L to maximize the matching measure. 5) Then, using the weighted overlap addition procedure, add at least one concealed frame from 2) and at least one frame from 3). This weighted overlap addition can be performed in a manner known to those of skill in the art, but may preferably be optimized for later disclosure herein. 6) The resulting at least one frame is used as an input to the phase fitting filtering described above starting at the determined value L. If L is greater than the threshold, activate some filters and sweep the coefficients at different moments and time intervals. In this case, the sum of the individual L values is the determined value L.
Effectively, for audio sampled at 8 or 16 kHz or residual audio, the above threshold may be chosen to be a value in the range 5-50. More effectively, in the case of vocalization or residual vocalization, testing of L_test concealed samples and their continuation to subsequent frames is achieved by cyclically shifting the sample in the first pitch period of the frame. To. This effectively allows the unnormalized correlation measure to correlate the full pitch period to be used as the matching measure in order to obtain the preferred cyclic shift L.
FIG. 9 shows an embodiment of such a method. In this figure, the phase adjustment produces a smooth transition between the signal frame 900 and subsequent frames. This is achieved as follows. That is, the concealed signal 910 is generated from the signal frame 900 and the frames before it. This concealment signal may be generated using the methods disclosed herein, or may be generated using other methods well known from state-of-the-art technology. The concealment signal is multiplied by window 920 and added to another window 930 by 925. Here, the window 930 is multiplied by the signal 940 generated as follows. That is, the concealment signal 940 uses from subsequent samples 950 and possibly 960, by effectively applying concealment methods such as those disclosed herein, or by using other methods well known from the latest technology. Generated by this and concatenated with subsequent sample 950. The number of samples in the concealment 940 is optimized to maximize the matching of the concealment 910 with the concatenation of the 940 and subsequent samples 950.
Effectively, the normalized correlation may be used as a measure of this matching. Further, to reduce computational complexity, matching may be limited to include one pitch period for vocalized or residual vocalized speech. In this case, the concealed sample 940 may be acquired as the first part of the one-pitch cycle cyclic shift, thus eliminating the need to normalize the one-pitch cycle correlation measure. This omits the calculation for calculating the normalization coefficient. With respect to the indexing and interpolation operations previously described in the detailed description of this preferred embodiment, the window is also effectively a function of the matching quality indicator and / or a function of frequency, effectively multistage. It may be implemented as a delay line. The calculation of the filter 970 is as follows. The first L samples resulting from the overlap addition procedure are sent directly to its output and used to set up the initial state of the filter. After that, the filter coefficients are initialized as described above, and as the filter filters the samples L + 1 to the destination, these coefficients gradually remove the delay of L samples as described above. Adjusted to.
Again, in the above procedure, the method of optimizing window weights by maximizing the matching criteria described above is applied to the frequency-dependent weights and matched filters of window functions in the form of multi-stage delay lines or other parametric filter formats. Generalization of is also applied. In an effective embodiment, the temporal expansion of the frequency-dependent filter weights is the following three overlapping addition sequences, namely the fade-down of at least one concealed frame from the first earlier frame, the second temporal. Fade up with these filtered versions of the filter and its subsequent fade down again to match hidden frames from the frame after being retrieved in reverse index order, after a third time. Achieved by a sequence consisting of a fade-up of at least one frame. In another effective set of embodiments, the temporal expansion of frequency-dependent filter weights is the following four overlapping addition sequences, namely the fade-down of at least one concealed frame from the first earlier frame, the second. Fade up with these filtered versions of the filter and then its re-fade down to match the hidden frames from the frame after it is retrieved in the reverse index order of time, the third of this Achieved by a sequence consisting of a time-later filtered version frame fade-up and its re-fade-down to further improve matching, and finally a fourth time-later at least one frame fade-up. Will be done. More effective embodiments of the weighted overlap addition method will be disclosed later herein.
For the smoothing and equalization operation 610 in the embodiment in which the residual region sample is used as part of the information representing the audio signal, the smoothing and equalization is effectively a comb filter or periodic notch. Pitch-adaptive filtering, such as a filter, may be used to apply to this residual signal. Further, effectively, Wiener or Kalman filtering may be applied using a long-term correlation filter with noise added as an unfiltered residual model. In this method of applying a Wiener or Kalman filter, the noise variance in the model is applied to adjust the degree of smoothing and equalization. This component has traditionally been applied in Weena and Kalman filtering theory to model the presence of unwanted noise components, which is a somewhat counterintuitive use. When this applies in this innovation, its purpose is to set the level of smoothing and equalization. In the context of this innovation, a third method effectively applies to smoothing and equalizing residual signals as an alternative to pitch-adaptive comb filters or notch filtering and Wiener or Kalman filtering. By this third method, either the sample amplitude effectively applied, for example to unvoiced speech, or the contiguous vector of the sample effectively, for example, applied to vocalized speech, is more and more similar. Will be made. The procedures that can achieve this are outlined below in relation to the vocal voice vector and the unvoiced voice sample, respectively.
For vocalized speech, continuous samples of speech or residue are collected in multiple vectors, where each vector is equal to one pitch period and has several samples. For convenience of explanation, this vector is represented by v (k) here. Next, in this method, the residual vector r (k) is subjected to the surrounding vectors v (k-L1), v (k-L1 + 1), ..., v (k-1) and v (k) by some means. Obtained as a component of v (k) that could not be found in +1), v (k + 2), ..., v (k + L2). For convenience of explanation, the components found in the surrounding vector are represented by a (k). The residual vector r (k) is subsequently manipulated to reduce its audibility in some linear or non-linear way, and at the same time inserts component a (k) into this manipulated version of r (k). The naturalness of the finally reconstructed vector achieved by re-doing is preserved.
This results in a smoothed and equalized form of uttered speech or utterance residual speech. The following is a simple embodiment of the above principle, using the matrix-vector notation for convenience and the concept of linear combination and least squares defining a (k) to simplify the example. Shown. However, this is merely an example of a simple and single embodiment of the general principles of smoothing and equalization described above.
For the purpose of this example, the matrix M (k) is defined as follows.
[Number 4] M (k) = [v (k-L1) v (k-L1 + 1) ... v (k-1) v (k + 1) v (k + 2) ... v (k + L2) ]
From the above equation, a (k) can be calculated as the least squares estimation of v (k) given, for example, M (k).
[Number 5] a (k) = M (k) inv (trans (M (k)) M (k)) v (k)
Here, inv () represents matrix inversion or pseudo-inversion, and trans () represents the transpose of the matrix. Therefore, the residual r (k) can be calculated by, for example, the following subtraction.
[Number 6] r (k) = v (k) -a (k)
An example of the operation of r (k) is, for example, that the maximum absolute value of the sample is at a level equal to the maximum amplitude of r (k) closest to the starting point of the previous or next concealment procedure, or at the same position in the vector. In order to limit the amplitude of the sample closest to the start point of the concealment procedure before and after in the vector multiplied by some coefficient, the peak of this vector is clipped and removed. The manipulated residual rm (k) is subsequently combined with the a (k) vector and reconstructed in an equalized form of v (k). Here, this is represented by ve (k) for convenience. As an example, this combination can be achieved by the following simple addition.
[Number 7] ve (k) = alpha * rm (k) + a (k)
The parameter alpha in this example may be set to 1.0 and effectively may be chosen to be less than 1.0, one of which is 0.8.
For unvoiced voice, another smoothing and equalization method may be effectively used. An example of unvoiced speech smoothing and equalization calculates a polynomial fitting with the amplitude of the residual signal in the logarithmic region. As an example, a quadratic polynomial and a log10 region may be used. After converting the polynomial fitting from the logarithmic region to the linear region and back, the fitting curve is normalized to 1.0 at the point corresponding to the starting point of the anteroposterior procedure. Subsequently, the fitting curve is limited downward, eg, 0.5, after which the amplitude of the residual signal may be divided by the fitting curve so as to smoothly equalize the amplitude deformation of the unvoiced residual signal.
With respect to the weighted overlap addition procedure, some, but not all, applications thereof have previously been disclosed herein of how to activate the input signals of indexing and interpolation operations 620 and phase adjustment filtering 970. .. These procedures may be performed in a manner well known to those of skill in the art. However, in a preferred embodiment of the weighted overlap addition procedure, the methods disclosed below may be effectively used.
In a simple embodiment of the weighted overlap addition procedure that is modified in response to the matching quality indicator, the first window is multiplied by the first subsequence and the second window is the second subsequence. It is assumed that they are multiplied and the product of these two is input to the overlap addition operation. Here, as an example, the first window is a tapered window such as a monotonically decreasing function, and the second window is a tapered window such as a monotonically increasing function. Second, to simplify the example, the second window is parameterized by the product of the basic window shape and the scalar multiplier. Here, target is defined as the first subsequence above, w_target is defined as the first subsequence for each sample multiplied by the tapered window, and w_regressor is the basic window shape of the tapered window. It is defined as the second subsequence for each multiplied sample, and coef is defined as the above scalar multiplier. The scalar multiplier component of the second window can now be optimized by minimizing the sum of the squared errors between the target and the result of the overlap addition operation. For convenience, using the matrix-vector notation, the above problem can be formulated as a minimization of the total squared difference between the target and the quantity shown by the following equation.
[Number 8] w_target + w_regressor * coef
From this, the vectors T and H are defined as follows.
[Number 9] T = target-w_target
[Number 10] H = w_regressor
The solution to this optimization problem is given by the following equation.
[Number 11] coef = inv (trans (H) * H) * trans (H) * T
Where inv () represents a scalar or matrix inversion, trans () represents a matrix or vector transpose, and * is a matrix multiplication or vector multiplication. Next, as a central element in the invention disclosed herein, this method may be extended to optimize the actual shape of the window. One way to achieve this is as follows. We define a set of shapes as a set to obtain the desired window as a linear combination of the elements contained in the set of shapes. Here we define H as each column of H being one shape from this set multiplied by the second subsequence above for each sample, and coefs these in the optimized window function. Defined as a column vector containing unknown weights of the shape. Using these definitions, the above equations that formulate the problem and its solution are in turn applied for a more general window-shaped solution. Of course, the roles of the first and second windows may be compatible in the above task, so the optimization execution target here is the first window.
A more advanced embodiment of the present invention optimizes both of these window shapes at the same time. This is probably the equivalent of the first set of window shapes and is effectively selected as the time-reversal indexing of the samples in each of the window shapes in the first set of window shapes, the basic window. This is done by defining a second set of shapes. Here, w_target is defined as a matrix in which each column is the basic window shape from the second set of window shapes multiplied by the first subsequence for each sample, and coef is defined as first, It is defined as a column vector containing the weights for the first window and secondly containing the weights for the second window. Now, the more general problem can be formulated as the minimization of the sum of squared differences between the target and the quantity shown by the following equation.
[Number 12] [w_target w_regressor] * coef
Here, square brackets [] are used to form a matrix from submatrixes or vectors. Next, from now on, the vectors T and H are defined as follows.
[Number 13] T = target
[Number 14] H = [w_target w_regressor]
The solution to this optimization is given by the following equation.
[Number 15] coef = inv (trans (H) * H) * trans (H) * T
Further, a more advanced embodiment of the present invention optimizes not only the instantaneous window shape, but also the window with the optimized frequency dependent weights. One embodiment of the present invention applies the form of a multi-stage delay line, but the present invention in general is not limited to this form in any case. One way to achieve this generalization is to replace each column with several columns, each of which performs a basic window shape multiplication for each sample, in the w_target and w_regressor definitions above. The basic window shape is the column that some of these columns replace, except that this basic window shape corresponds to a specific position on the multi-stage delay line for each sample at that temporal position. Corresponds to the column where the sequence is multiplied.
Effectively, the optimization of the coefficients in these methods takes into account the weights, constraints or sequential calculations of the coefficients without departing from the inventions disclosed herein. Such weights effectively include weights that tend to give greater weight to the coefficients corresponding to the lower absolute delay values. Such sequential calculations effectively use low absolute delay value coefficients, first using only these coefficients to minimize the sum of squared errors, and then this process, with respect to increasing delay values. Calculations may be made to iterate only for the errors remaining from the early steps of this process.
In general, embodiments of the present invention employ several subsequences as optimization goals. Generally speaking, the optimization minimizes the distortion function, which is a function of the subsequences of these goals and the output from the weighted overlap addition system. This optimization may impose various constraints on the overall weight selection and delay and overlap addition without departing from the present invention. Depending on the exact choice of shape, the effect of the overlap addition is effectively faded out of the subsequence following the overlap addition region in time.
FIG. 10 shows an embodiment of the disclosed overlap addition method. The present invention is not limited to the exact structure in this figure, and thus this figure is merely to illustrate one embodiment of the present invention. In FIG. 10, one subsequence 1000 is input with another subsequence 1010 in a time and frequency shape optimized overlap addition. Each of these subsequences is input to a separate delay line. In this figure, z indicates the time lead for one sample, and z-1 indicates the time delay for one sample. The delays of 1, -1 and 0 selected are purely exemplary and more or less other delays can be effectively used in the present invention. Each delayed version of each subsequence is then multiplied by some basic window shape, and each of these results is multiplied by a factor that should be found at the same time as the other coefficients during the optimization process. .. After multiplication by these coefficients, the resulting subsequences are added to yield an output of 1020 from a time and frequency shape optimized overlap addition. Coefficient optimization 1030 takes subsequences 1040 and 1050 as inputs in the example in FIG. 10 and minimizes the distortion function, which is a function of 1040 and 1050 and output 1020.
The quotation marks indicating the drawings in the claims are described solely for the purpose of clarity. These quotation marks, which refer to exemplary embodiments in the figures, should not be construed as limiting the scope of the claims in any case.
<figref num="1">FIG. 6 is a block diagram showing a known end-to-end packet-switched voice transmission system affected by loss, delay, delay jitter and / or clock skew.</figref><figref num="2">An exemplary receiver subsystem that achieves jitter buffering, decoding and concealment, and playback output buffering under the control of a control unit is shown.</figref><figref num="3">It is a block diagram which shows the relay subsystem of the packet switching channel which is affected by clock skew, loss, delay and delay jitter.</figref><figref num="4">An exemplary relay subsystem that achieves input buffering, output buffering and, if necessary, transcoding and concealment under the control of a control unit is shown.</figref><figref num="5">It is a block diagram which shows a series of preferable embodiments of this invention.</figref><figref num="5A">It is a sketch depicting a subsequence in a concealed frame, and the starting point of the frame is a subsequence based on the last buffered subsequence in the reverse order of time.</figref><figref num="5B">To show another example with a larger sequence of subsequences in a hidden frame, the starting point of the frame is the last two buffered subsequences in reverse order of time, and successive subsequences are reverse time. Based on subsequences further buffered in.</figref><figref num="5C">Shows the sample count index in the index pattern formatted by step back and read length.</figref><figref num="6">A sketch depicting signals related to indexing and interpolation functions.</figref><figref num="7">It is a flowchart which shows one method which can execute the decision logic of a stop criterion.</figref><figref num="8">It is a flowchart which shows one method which can achieve the smoothing and equalization, the stop criterion and the iterative simultaneous optimization of the permissible number of iterations.</figref><figref num="9">The use of cyclic shift and overlap addition associated with the initialization and supply of phase adjustment filters is shown.</figref><figref num="10">An embodiment of the disclosed weighted overlap addition procedure is shown.</figref>
13 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2014038347A | Cited by | Japan | Examiner |
| US9270722B2 | Cited by | United States of America | Applicant |
| JP2005315973A | Cites | Japan | – |
| JP2002542521A | Cites | Japan | – |
| JP2003316670A | Cites | Japan | – |
| WO02095731A1 | Cites | World Intellectual Property Organization (WIPO) | – |
| JP2004501391A | Cites | Japan | – |
| WO02071389A1 | Cites | World Intellectual Property Organization (WIPO) | – |
| JP2002328691A | Cites | Japan | – |
| JP200477961A | Cites | Japan | – |
| JP10209977A | Cites | Japan | – |
78 members in 15 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| PA200500146 | Denmark | A | |
| PA200500146 | Denmark | A | |
| PA200500146 | Denmark | – | |
| 2006000053 | Denmark | W | |
| 2006000053 | Denmark | W | |
| 2005200500146 | – | – | – |
| 2006000053 | – | – | – |
| DK2005PA00146 | – | – | – |
| DKPA200500146 | – | – | – |
| WO2006DK00053 | – | – | – |
Members78
| Document | Office | Kind | |
|---|---|---|---|
| AU2006208528A1 | Australia | A1 | |
| AU2006208529A1 | Australia | A1 | |
| AU2006208530A1 | Australia | A1 | |
| CA2596337A1 | Canada | A1 | |
| CA2596338A1 | Canada | A1 | |
| CA2596341A1 | Canada | A1 | |
| WO2006079348A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2006079349A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2006079350A1 | World Intellectual Property Organization (WIPO) | A1 | |
| NO20074418L | Norway | L | |
| NO20074349L | Norway | L | |
| NO20074348L | Norway | L | |
| EP1846920A1 | European Patent Office (EPO) | A1 | |
| EP1846921A1 | European Patent Office (EPO) | A1 | |
| EP1849156A1 | European Patent Office (EPO) | A1 | |
| IL184864D0 | Israel | D0 | |
| IL184927D0 | Israel | D0 | |
| IL184948D0 | Israel | D0 | |
| KR20080001708A | Republic of Korea | A | |
| KR20080002756A | Republic of Korea | A | |
| KR20080002757A | Republic of Korea | A | |
| CN101120398A | China | A | |
| CN101120399A | China | A | |
| CN101120400A | China | A | |
| HK1108760A1 | Hong Kong, China | A1 | |
| ZA200706307B | South Africa | B | |
| US2008154584A1 | United States of America | A1 | |
| ZA200706534B | South Africa | B | |
| JP2008529072A | Japan | A | |
| JP2008529073A | Japan | A | |
| JP2008529074A | Japan | A | |
| US2008275580A1 | United States of America | A1 | |
| RU2007132728A | Russian Federation | A | |
| RU2007132729A | Russian Federation | A | |
| RU2007132735A | Russian Federation | A | |
| ZA200706261B | South Africa | B | |
| BRPI0607246A2 | Brazil | A2 | |
| BRPI0607247A2 | Brazil | A2 | |
| US2010161086A1 | United States of America | A1 | |
| AU2006208529B2 | Australia | B2 | |
| AU2006208530B2 | Australia | B2 | |
| RU2405217C2 | Russian Federation | C2 | |
| RU2407071C2 | Russian Federation | C2 | |
| IL184864A | Israel | A | |
| RU2417457C2 | Russian Federation | C2 | |
| CN101120399B | China | B | |
| AU2006208528B2 | Australia | B2 | |
| US8068926B2 | United States of America | B2 | |
| AU2006208528C1 | Australia | C1 | |
| CN101120398B | China | B | |
| US2012158163A1 | United States of America | A1 | |
| IL184948A | Israel | A | |
| EP1849156B1 | European Patent Office (EPO) | B1 | |
| KR101203244B1 | Republic of Korea | B1 | |
| KR101203348B1 | Republic of Korea | B1 | |
| KR101237546B1 | Republic of Korea | B1 | |
| CN101120400B | China | B | |
| JP5202960B2 | Japan | B2 | |
| CA2596341C | Canada | C | |
| JP5420175B2This record | Japan | B2 | |
| JP2014038347A | Japan | A | |
| CA2596338C | Canada | C | |
| CA2596337C | Canada | C | |
| US8918196B2 | United States of America | B2 | |
| US9047860B2 | United States of America | B2 | |
| US2015207842A1 | United States of America | A1 | |
| US9270722B2 | United States of America | B2 | |
| JP5925742B2 | Japan | B2 | |
| IL184927A | Israel | A | |
| NO338702B1 | Norway | B1 | |
| NO338798B1 | Norway | B1 | |
| EP1846920B1 | European Patent Office (EPO) | B1 | |
| BRPI0607251A2 | Brazil | A2 | |
| NO340871B1 | Norway | B1 | |
| ES2625952T3 | Spain | T3 | |
| EP1846921B1 | European Patent Office (EPO) | B1 | |
| BRPI0607247B1 | Brazil | B1 | |
| BRPI0607246B1 | Brazil | B1 |
33 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of appointment of power of attorneyJAPANESE INTERMEDIATE CODE: A7423RD03 | RD03 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| Notification of change in applicantJAPANESE INTERMEDIATE CODE: A711A711 | A711 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 |
Numbers
- Publication
- 5420175
- Publication, DOCDB
- 5420175
- Publication, EPODOC
- JP5420175B
- Application
- 2007552505
- Application, DOCDB
- 2007552505
- Application, EPODOC
- JP20070552505
Titles2
- Japanese
- 通信システムにおける隠蔽フレームの生成方法
- English
- How to generate a hidden frame in a communication system
Classification
- CPC, 7
- G10L19/005
- H04L65/764
- H04L65/00
- H04M3/18
- H04L12/28
- G10L19/24
- G10L19/26
- IPC, 3
- G10L19 005
- H04L47 43
- H04L49 9023
