Motion estimation techniques for video encoding
Abstract
Describes a video coding technique. In one example, a video encoding technique includes identifying a pixel location associated with a video block in a search space based on a motion vector associated with a group of video blocks in a video frame to be encoded, wherein The video block is spatially located at a certain position relative to the current video block of the video frame to be encoded. Then, an initial motion estimation routine can be performed for the current video block at the identified pixel position. By identifying a pixel position associated with a video block in the search space based on a motion vector associated with a group of video blocks in a video frame, it is easier to take advantage of spatial redundancy to speed up and improve the encoding process.

Term
Term ended
Projected expiry passed 18 June 2023, 3.3 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
54 claims: 12 independent, 42 dependent
- 1一种设备,其特征在于,它包括:一编码器,它使用一编码例程来编码视频帧,所述编码例程包括基于与一视频帧内的一组视频块相关联的计算运动矢量来标识一与搜索空间内的一视频块相关联的像素位置,所述组中的视频块位于相对于要编码的所述视频帧的当前视频块的确定位置上,以及对在所标识的像素位置上的所述当前视频块初始化一运动估计例程;以及一发送器,它发送所编码的视频帧。
- 2如权利要求1所述的设备,其特征在于,标识所述像素位置包括基于与位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块相关联的运动矢量来计算一组像素坐标。
- 3如权利要求2所述的设备,其特征在于,计算所述像素坐标组包括基于位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块的运动矢量来计算一中值。
- 4如权利要求2所述的设备,其特征在于,计算所述像素坐标组包括基于位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块的运动矢量来计算一平均值。
- 5如权利要求2所述的设备,其特征在于,计算所述像素坐标组包括基于位于相对于所述当前视频块的确定位置的所述视频块组中的视频块的运动矢量来计算一加权函数,其中,在所述加权函数中,向与空间上更邻近所述当前视频块的所述视频组中的视频块相关联的运动矢量给予比与空间上更远离所述当前视频块的所述视频块组中的视频块相关联的运动矢量更大的权值。
- 6如权利要求1所述的设备,其特征在于,所述编码器使用所述运动估计例程来编码所述当前视频块,所述运动估计例程包括:在所标识的像素位置周围定义一半径为(R)的圆;将所述当前块与同所述圆内的像素位置相关联的搜索空间的视频块相比较;标识所述圆内标识产生最低差值的视频块的像素位置;以及当所述圆内标识产生最低差值的视频块的所标识的像素位置对应于所述圆的圆心时,使用由所述圆内标识产生最低差值的视频块的所标识的像素位置定义的运动矢量来编码所述当前视频块。
- 7如权利要求1所述的设备,其特征在于,所述编码器使用所述运动估计例程来编码所述当前视频块,所述运动估计例程包括基于可用于编码所述当前视频帧的确定的计算资源量动态地调整要执行的计算数量。
- 8一种设备,其特征在于,它包括:一编码器,它通过以下步骤编码视频帧:为编码视频帧的当前视频块,在搜索空间内的一像素位置上初始化一运动估计例程;在所述像素位置周围定义一半径为(R)的圆;将所述当前视频块与同所述圆内的一组像素位置相关联的搜索空间的视频块相比较;标识所述圆内标识产生最低差值的视频块的像素位置;以及当所标识的像素位置对应于所述圆的圆心时,使用由所标识的像素位置所定义的运动矢量来编码所述当前视频块;以及一发送器,它发送所编码的视频帧。
- 9如权利要求8所述的设备,其特征在于,当所述半径为(R)的圆内标识产生最低差值的视频块的所标识的像素位置不对应于所述圆的圆心时,所述编码器通过以下步骤编码所述视频帧:在不对应于所述半径为(R)的圆的圆心的、标识产生最低差值的视频块的所述像素位置周围定义另一半径为(R′)的圆;标识所述半径为(R′)的圆内标识产生最低差值的视频块的像素位置;当所述半径为(R′)的圆内标识产生最低差值的视频块的所述像素位置对应于所述半径为(R′)的圆的圆心时,使用所述半径为(R′)的圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述当前视频块。
- 10如权利要求9所述的设备,其特征在于,R近似等于R′。
- 11如权利要求9所述的设备,其特征在于,当所述半径为(R)的圆内标识产生最低差值的视频块的所述像素位置不对应于所述半径为(R′)的圆的圆心时,所述编码器通过以下步骤编码所述视频帧:在不对应于所述半径为(R′)的圆的圆心的、标识产生最低差值的视频块的所述像素位置周围定义另一半径为(R″)的圆;标识所述半径为(R″)的圆内标识最低差值的视频块的像素位置;以及当所述半径为(R″)的圆内标识产生最低差值的视频块的所述像素位置对应于所述半径为(R″)的圆的圆心时,使用由所述半径为(R″)的圆内标识产生最低差值的视频块的所述像素位置所定义的运动矢量来编码所述当前视频帧。
- 12如权利要求8所述的设备,其特征在于,所述设备选自以下组:数字电视机、无线通信设备、个人数字助理、膝上计算机、台式机、数码相机、数码记录设备、具有视频能力的蜂窝无线电话、以及具有视频能力的卫星无线电话。
- 13一种视频编码方法,其特征在于,它包括:基于一视频帧内一组视频块的运动矢量来标识搜索空间内的一像素位置,所述组中的视频块位于相对于要编码的所述视频帧的当前视频块的确定位置上;以及在所标识的像素位置上为所述当前视频块来初始化一运动估计例程。
- 14如权利要求13所述的方法,其特征在于,位于相对于所述当前视频块的确定位置上的所述组中的所述视频块包括与所述当前视频块相邻的视频块。
- 15如权利要求13所述的方法,其特征在于,标识所述像素位置包括基于为位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块所计算的运动矢量来计算一组像素坐标。
- 16如权利要求15所述的方法,其特征在于,计算所述像素坐标组包括基于为位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块计算的运动矢量来计算一中值。
- 17如权利要求15所述的方法,其特征在于,计算所述像素坐标组包括基于为位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块计算的运动矢量来计算一平均值。
- 18如权利要求15所述的方法,其特征在于,计算所述像素坐标组包括基于为位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块计算的运动矢量来计算一加权函数,其中,在所述加权函数中,对为空间上更邻近所述当前视频块的所述组中的视频块计算的运动矢量给予比为空间上更远离所述当前视频块的所述组中的视频块计算的运动矢量更大的权值。
- 19如权利要求13所述的方法,其特征在于,它还包括使用所述运动估计例程来编码所述当前视频帧。
- 20如权利要求29所述的方法,其特征在于,所述运动估计例程包括:在所标识的像素位置周围定义一半径为(R)的圆;将所述当前视频块与同所述圆内的像素位置相关联的搜索空间的视频块相比较;标识所述圆内标识产生最低差值的视频块的像素位置;以及当所述圆内标识产生最低差值的视频块的所标识的像素位置对应于所述圆的圆心时,使用由所述圆内标识产生最低差值的视频块的所标识的像素位置定义的运动矢量来编码所述当前视频块。
- 21如权利要求20所述的方法,其特征在于,当所述圆内标识产生最低差值的所标识的像素位置不对应于所述圆的圆心时:在所述圆内标识产生最低差值的视频块的所标识的像素位置周围定义另一半径为(R′)的圆;标识所述半径为(R′)的圆内标识产生最低差值的视频块的像素位置;以及当所述半径为(R′)的圆内标识产生最低差值的视频块的所标识的像素位置对应于所述半径为(R′)的圆的圆心时,使用由所述半径为(R′)的圆内标识产生最低差值的视频块的所标识的像素位置定义的运动矢量来编码所述当前视频块。
- 22如权利要求19所述的方法,其特征在于,所述运动估计例程包括基于可用于编码所述当前视频块的确定的计算资源量动态地调整要执行的计算数量。
- 23一种方法,其特征在于,它包括:为编码视频帧的当前视频块,在搜索空间内的一像素位置上初始化一运动估计例程;在所述像素位置周围定义一半径为(R)的圆;将所述当前视频块与同所述圆内的一组像素位置相关联的搜索空间的视频块相比较;标识所述圆内标识产生最低差值的视频块的像素位置;以及当所标识的像素位置对应于所述圆的圆心时,使用由所标识的像素位置定义的运动矢量来编码所述当前视频块。
- 24如权利要求23所述的方法,其特征在于,它还包括当所标识的像素位置不对应于所述圆的圆心时:在所标识的像素位置周围定义另一半径为(R′)的圆;标识所述半径为(R′)的圆内标识产生最低差值的视频块的像素位置;以及当所述半径为(R′)的圆内标识产生最低差值的视频块的所标识的像素位置对应于所述半径为(R′)的圆的圆心时,使用由所述半径为(R′)的圆内标识产生最低差值的视频块的所标识的像素位置定义的运动矢量来编码所述当前视频帧。
- 25如权利要求24所述的方法,其特征在于,R近似等于R′。
- 26如权利要求24所述的方法,其特征在于,它还包括当所述半径为(R′)的圆内标识产生最低差值的视频块的所标识的像素位置不对应于所述半径为(R′)的圆的圆心时:在所述半径为(R′)的圆内标识产生最低差值的视频块的所标识的像素位置周围定义另一半径为(R″)的圆;标识所述半径为(R″)的圆内标识产生最低差值的视频块的像素位置;以及当所述半径为(R″)的圆内标识产生最低差值的视频块的所标识的像素位置对应于所述半径为(R″)的圆的圆心时,使用由所述半径为(R″)的圆内标识产生最低差值的视频块的所标识的像素位置定义的运动矢量来编码所述当前视频块。
- 27一种装置,其特征在于,它包括:一存储器,它储存计算机可执行指令;以及一处理器,它执行所述指令,以便:基于与一视频帧内一组视频块相关联的计算的运动矢量来标识与搜索空间内一视频块相关联的一像素位置,所述组中的视频块位于相对于要编码的当前视频块的确定的位置上;以及在所标识的像素位置上为所述当前视频块初始化一运动估计例程。
- 28如权利要求27所述的装置,其特征在于,所述处理器基于位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块相关联的运动矢量来计算一中值,用于标识所述像素位置。
- 29如权利要求27所述的装置,其特征在于,所述处理器基于位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块的像素坐标来计算一平均值,用于标识所述像素位置。
- 30如权利要求27所述的装置,其特征在于,所述处理器基于位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块的运动矢量来计算一加权函数,用于标识所述像素位置,其中,在所述加权函数中,向所述组中与空间上更邻近所述当前视频块的视频块相关联的运动矢量给予比所述组中与空间上更远离所述当前视频块的视频块相关联的运动矢量更大的权值。
- 31如权利要求27所述的装置,其特征在于,所述处理器执行所述指令以执行所述运动估计例程,其中,所述运动估计例程包括:在所标识的像素位置周围定义一半径为(R)的圆;将所述当前视频块与同所述圆内的像素位置相关联的搜索空间的视频块相比较;标识所述圆内标识产生最低差值的视频块的像素位置;以及当所述圆内标识产生最低差值的视频块的所述像素位置对应于所述圆的圆心时,使用由所述圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述当前视频块。
- 32如权利要求27所述的装置,其特征在于,所述处理器执行所述指令以执行所述运动估计例程,其中,所述运动估计例程包括基于可用于编码所述当前视频块的确定的计算资源量动态地调整要执行的计算的数量。
- 33一种装置,其特征在于,它包括:一存储器,它储存计算机可执行指令;以及一处理器,它执行所述指令,以便:为编码一视频帧的当前视频块,在搜索空间内的一像素位置上初始化一运动估计例程;在所述像素位置周围定义一半径为(R)的圆;将所述当前块与同所述圆内一组像素位置相关联的搜索空间的视频块相比较;标识所述圆内标识产生最低差值的视频块的像素位置;以及当所述圆内标识产生最低差值的视频块的所述像素位置对应于所述圆的圆心时,使用由所述圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述当前视频块。
- 34如权利要求33所述的装置,其特征在于,当所述半径为(R)的圆内标识产生最低差值的视频块的所述像素位置不对应于所述圆的圆心时,所述处理器执行指令,以便:在不对应于所述半径为(R)的圆心的、标识产生最低差值的视频块的所述像素位置周围定义另一半径为(R′)的圆;标识所述半径为(R′)的圆内标识产生最低差值的视频块的像素位置;以及当所述半径为(R′)的圆内标识产生最低差值的视频块的所述像素位置对应于所述半径为(R′)的圆的圆心时,使用由所述半径为(R′)的圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述当前视频块。
- 35一种依照MPEG-4标准编码视频块的装置,其特征在于,所述装置:基于与一视频帧内一组视频块相关联的计算的运动矢量来标识与搜索空间内一视频块相关联的一像素位置,所述视频块组位于相对于要编码的所述视频帧的当前视频块的确定位置上;以及在所标识的像素位置上为所述当前视频块初始化一运动估计例程。
- 36如权利要求35所述的装置,其特征在于,所述装置包括一数字信号处理器,它执行计算机可读指令以依照MPEG-4标准编码所述视频块。
- 37一种依照MPEG-4标准编码视频块的装置,其特征在于,所述装置:为编码一视频帧的当前视频块,在搜索空间中的一像素位置上初始化一运动估计例程;在所述像素位置周围定义一半径为(R)的圆;将所述当前视频块与同所述圆内一组像素位置相关联的搜索空间的视频块相比较;标识所述圆内标识产生最低差值的视频块的像素位置;以及当所述圆内标识产生最低差值的视频块的所述像素位置对应于所述圆的圆心时,使用由所述圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述当前视频块。
- 38如权利要求37所述的装置,其特征在于,所述装置包括一数字信号处理器,它执行计算机可读指令以依照MPEG-4标准编码所述视频块。
- 39一种包括指令的计算机可读媒质,其特征在于,当在编码符合MPEG-4标准的视频序列的设备中执行所述指令时:基于与一视频帧内一组视频块相关联的计算的运动矢量来标识与搜索空间内一视频块相关联的一像素位置,所述组中的视频块位于相对于要编码的所述视频帧的当前视频块的确定位置上;以及在所标识的像素位置上为所述当前视频块初始化一运动估计例程。
- 40如权利要求29所述的计算机可读媒质,其特征在于,位于相对于所述当前视频块的确定位置上的所述视频块包括与所述当前视频块相邻的视频块。
- 41如权利要求39所述的计算机可读媒质,其特征在于,它还包括指令,当所述指令被执行时,通过基于与位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块相关联的运动矢量来计算一组像素坐标,用于标识所述像素位置。
- 42如权利要求41所述的计算机可读媒质,其特征在于,计算所述像素坐标组包括基于位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块的运动矢量来计算一中值。
- 43如权利要求41所述的计算机可读媒质,其特征在于,计算所述像素坐标组包括基于位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块的运动矢量计算一平均值。
- 44如权利要求41所述的计算机可读媒质,其特征在于,计算所述像素坐标组包括基于与位于相对于所述当前视频块的确定位置上的所述视频块组中的视频块相关联的运动矢量来计算一加权函数,其中,在所述加权函数中,对与空间上更邻近所述当前视频块的所述组中的视频块相关联的运动矢量给予比空间上更远离所述当前视频块的所述组中的视频块相关联的运动矢量更大的权值。
- 45如权利要求29所述的计算机可读媒质,其特征在于,它还包括指令,当所述指令被执行时,使用所述运动估计例程来编码所述当前视频块。
- 46如权利要求45所述的计算机可读媒质,其特征在于,所述运动估计例程包括:在所标识的像素位置周围定义一半径为(R)的圆;将所述当前视频块与同所述圆内的像素位置相关联的搜索空间的视频块相比较;标识所述圆内标识产生最低差值的视频块的像素位置;以及当所述圆内标识产生最低差值的视频块的所述像素位置对应于所述圆的圆心时,使用由所述圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述当前视频块。
- 47如权利要求46所述的计算机可读媒质,其特征在于,所述运动估计例程还包括当所述半径为(R)的圆内标识产生最低差值的视频块的所述像素位置不对应于所述圆的圆心时:在不对应于所述半径为(R)的圆的圆心的、标识产生最低差值的视频块的所述像素位置周围定义另一半径为(R′)的圆;标识所述半径为(R′)的圆内标识产生最低差值的视频块的像素位置;以及当所述半径为(R′)的圆内标识产生最低差值的视频块的所述像素位置对应于所述半径为(R′)的圆的圆心时,使用所述半径为(R′)的圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述当前视频帧。
- 48如权利要求45所述的计算机可读媒质,其特征在于,所述运动估计例程包括基于可用于编码所述当前视频块的确定的计算资源量动态地调整要执行的计算数量。
- 49一种包括指令的计算机可读媒质,其特征在于,当在编码符合MPEG-4的视频序列的设备中执行所述指令时:为编码一视频帧的当前视频块,在搜索空间内的一像素位置上初始化一运动估计例程;在所述像素位置周围定义一半径为(R)的圆;将所述当前视频块与同所述圆内的一组像素位置相关联的搜索空间的视频块相比较;标识所述圆内标识产生最低差值的视频块的像素位置;以及当所述圆内标识产生最低差值的视频块的所述像素位置对应于所述圆的圆心时,使用由所述圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述当前视频块。
- 50如权利要求49所述的计算机可读媒质,其特征在于,当所述半径为(R)的圆内标识产生最低差值的视频块的所述像素位置不对应于所述圆的圆心时,当所述指令被执行时:在不对应于所述半径为(R)的圆的圆心的、标识产生最低差值的视频块的所述像素位置周围定义另一半径为(R′)的圆;标识所述半径为(R′)的圆内标识产生最低差值的视频块的像素位置;以及当所述半径为(R′)的圆内标识产生最低差值的视频块的所述像素位置对应于所述半径为(R′)的圆的圆心时,使用由所述半径为(R′)的圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述当前视频块。
- 51如权利要求50所述的计算机可读媒质,其特征在于,R近似等于R′。
- 52如权利要求50所述的计算机可读媒质,其特征在于,当所述半径为(R′)的圆内标识产生最低差值的视频块的所述像素位置不对应于所述半径为(R′)的圆的圆心时,当所述指令被执行时:在不对应于所述半径为(R′)的圆的圆心的、标识产生最低差值的视频块的所述像素位置的周围定义另一半径为(R″)的圆;标识所述半径为(R″)的圆内标识产生最低差值的视频块的像素位置;当所述半径为(R″)的圆内标识产生最低差值的视频块的所述像素位置对应于所述半径为(R″)的圆的圆心时,使用由所述半径为(R″)的圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述的当前视频帧。
- 53一种装置,其特征在于,它包括:用于基于与一视频帧内一组视频块相关联的计算的运动矢量来标识与搜索空间内一视频块相关联的一像素位置的装置,所述组中的视频块位于相对于要编码的所述视频帧的当前视频块的确定位置上;以及用于在所标识的像素位置上为所述当前视频块初始化一运动估计例程的装置。
- 54一种装置,其特征在于,它包括:用于为编码一视频帧的当前帧,在搜索空间内一像素位置上初始化一运动估计例程的装置;用于在所述像素位置周围定义一半径为(R)的圆的装置;用于将所述当前视频块与同所述圆内的一组像素位置相关联的搜索空间的视频块相比较的装置;用于标识所述圆内标识产生最低差值的视频块的像素位置的装置;以及用于当所述圆内标识产生最低差值的视频块的所述像素位置对应于所述圆的圆心时,使用由所述圆内标识产生最低差值的视频块的所述像素位置定义的运动矢量来编码所述当前视频帧的装置。
Independent claims54
75 paragraphs, as filed
Motion estimation technology for video coding
Technical field
The present disclosure relates to digital video processing, and in particular to the encoding of video sequences.
Background Digital video capabilities can be incorporated into a variety of devices, including digital televisions, digital direct broadcasting systems, wireless communication devices, personal digital assistants (PDAs), laptop computers, desktop computers, digital cameras, digital recording devices, Cellular or satellite radio phones, etc. These and other digital video devices provide significant improvements over conventional analog video systems in creating, modifying, sending, storing, recording, and playing full-motion video sequences.
Many different video coding standards have been established for coding digital video sequences. For example, the Moving Picture Experts Group (MPEG) has developed several standards, including MPEG-1, MPEG-2, and MPEG-4. Other standards include ITU H.263, QuickTimeTM technology developed by Apple Inc. of Cupertino, California, Video for WindowsTM developed by Microsoft Corporation of Redmond, Washington, IndeoTM developed by Intel Corporation, Washington State RealVideoTM from RealNetworks in Seattle and CinepakTM developed by SuperMac. These and other standards, including those that will be developed, will continue to evolve.
Many video coding standards achieve increased transmission rates by encoding data in a compressed manner. Compression can reduce the total amount of data that needs to be sent for effective transmission of video frames. For example, the MPEG standard uses graphics and video compression technology that is designed to conveniently transmit video and images on a bandwidth narrower than that available without compression. Specifically, the MPEG standard supports a video coding technique that uses the similarity between consecutive video frames, called temporal or inter-frame correlation, to provide inter-frame compression. The inter-frame compression technology takes advantage of the data redundancy between frames by converting the pixel-based representation of the video frame into a motion representation. In addition, video coding techniques can use similarity within frames, called spatial or intra-frame correlation, to further compress video frames. Intra-frame compression is usually based on a texture coding process used to compress still images, such as discrete cosine transform (DCT) coding.
To support compression technology, many digital video devices include encoders for compressing digital video sequences and decoders for decompressing digital video sequences. In many cases, the encoder and decoder form an integrated encoder/decoder (CODEC (CODEC)) that operates on the intra-frame pixel blocks that define a sequence of video images. For example, in the MPEG-4 standard, the encoder of the transmitting device usually divides the video frame to be transmitted into macroblocks containing smaller image blocks. For each macro block in the video frame, the encoder searches for the immediately preceding video frame in the macro block to identify the most similar macro block, encodes the difference between the macro blocks to be sent, and instructs the use of The motion vector of which macroblock of the previous frame was coded. The decoder of the receiving device receives the difference between the motion vector and the code, and performs motion compensation to generate a video sequence.
The video encoding process is computationally intensive, especially when using motion estimation. For example, the process of comparing the video block to be encoded with the video block of the previously transmitted frame requires a lot of calculation. Therefore, improved encoding techniques are highly desired, especially for wireless devices or other portable video devices where computing resources are more limited and power consumption is very important. At the same time, improved compression is desired to reduce the bandwidth required for effective transmission of video sequences. Improvements in one or more of these factors can facilitate or improve the real-time encoding of video sequences, especially in wireless and other limited bandwidth settings.
Overview This invention describes video encoding techniques that can be used to encode video sequences. For example, video encoder technology may involve identifying pixel locations in a search space that are associated with video blocks based on the motion vectors of a set of video blocks within a video frame to be encoded. The video blocks in the group of video blocks may include video blocks located at positions determined relative to the current video block of the video frame to be encoded. Then a motion estimation routine can be initialized for the current video block at the identified pixel position. By identifying the pixel position associated with the video block in the search space based on the calculated motion vector associated with a set of video blocks within the video frame, it is easier to use spatial redundancy to speed up and improve the encoding process . In various examples, the initialized position may be calculated using a linear or non-linear function based on the motion vector of a group of video blocks at a determined position relative to the video block to be encoded. For example, a median function, an average function, or a weighting function based on the motion vector of the group of video blocks can be used.
After performing the initialization motion estimation routine of the current video block of the encoded video frame at the pixel position in the search space, the motion estimation routine of encoding the current video block can be executed. To reduce the number of calculations in the encoding process, the motion estimation routine may include a non-exhaustive search of video blocks in the search space. For example, a motion estimation routine may include defining a circle with a radius (R) around the initialized pixel position, and comparing the current video block to be encoded with a video block in the search space associated with a set of pixel positions within the circle . The circle may be large enough to include at least five pixels. However, by defining the circle to be large enough to include at least nine pixels, such as a central pixel and eight pixels surrounding the central pixel, the process can be improved by anticipating video motion in each direction. Larger or smaller circles can also be used. The routine may also include identifying the pixel position of the video block producing the lowest difference in the circle, and when the pixel position corresponds to the center of the circle, using the motion defined by the pixel position of the video block producing the lowest difference in the circle. Vector to encode the current video block.
However, if the pixel location of the video block that produces the lowest difference in the circle does not correspond to the center of the circle with radius (R), then a new pixel location can be defined around the pixel location of the video block that produces the lowest difference in the circle. The radius of the circle is (R'). In this case, the routine may also include determining the pixel position of the video block that produces the lowest difference in the circle with radius (R), and when the pixel position corresponds to the circle with radius (R). At the center of the circle, use the motion vector defined by the pixel to encode the current video block. The routine can continue in a similar manner, defining additional circles as needed, until the video block that produces the lowest difference corresponds to the center of the circle.
These and other technologies described in more detail below can be implemented in digital video equipment in hardware, software, firmware, or any combination thereof. If implemented in software, the technology is aimed at a computer-readable medium including program code, and when the program code is executed, one or more of the coding technologies described in the present invention can be executed. Additional details of various embodiments are set forth in the appended claims and described later. Reading the following description, drawings and claims, other features, purposes and advantages can be clear.
BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1 shows a block diagram of an example system in which a source device sends an encoded video data sequence to a sink device.
Figure 2 shows a block diagram of an example digital video device incorporating a video encoder that encodes a digital video sequence.
Fig. 3 is a conceptual illustration of an example macroblock of video data.
Figure 4 is a conceptual illustration of an example search space.
Figure 5 shows a flowchart of a video coding technique that can be implemented in a digital video device to initialize a motion estimation routine.
6-8 are conceptual illustrations of exemplary video frames in which the current video block is encoded using a technique similar to that shown in FIG. 5.
Figure 9 shows a flowchart of a video coding technique that can be implemented in a digital video device to perform motion estimation.
10-13 are conceptual illustrations of a search space in which a video coding technique similar to that shown in FIG. 9 is performed to find acceptable motion vectors.
Figure 14 shows a flowchart of a video coding technique that can be implemented in digital video equipment to improve real-time video coding.
Detailed Description Generally speaking, the present disclosure is directed to coding techniques that can be used to improve the coding of digital video data. These techniques can be performed by the encoder of digital video equipment to reduce the amount of calculation in some cases, speed up processing time, improve data compression, and possibly reduce power consumption during video encoding. In this way, these coding techniques can improve video coding in accordance with standards such as MPEG-4, and facilitate the realization of video coding in wireless devices where computing resources are more limited and power consumption is very important. In addition, these technologies can be configured to maintain interoperability with decoding standards such as the MPEG-4 decoding standard.
The video coding technology described in the present invention can realize an initialization technology that utilizes the spatial redundancy phenomenon. Spatial redundancy generally predicts that the video motion of a given video block may be similar to the video motion of another video block that is spatially adjacent to the given video block. This initialization technique can more easily take advantage of this phenomenon to perform initialization motion estimation at positions in the search space that have a very high probability of including video blocks that can be used for effective video coding.
More specifically, the initialization technique may use the motion vector calculated for the video block spatially adjacent to the video block to be encoded to identify the position in the search space where the motion estimation routine can be initialized, that is, the motion estimation routine in the search space The starting pixel position. For example, as described in more detail later, the average pixel position, the median pixel position, or the value calculated using a weighting function can be calculated based on a previously determined motion vector for a video block spatially adjacent to the current video block to be encoded. Pixel position. Other linear or non-linear functions can also be used. In either case, by initializing the motion estimation routine in this way, video encoding can be accelerated by reducing the number of comparisons required to find a video block that is an acceptable match to the video block to be encoded in the search space. The motion estimation routine may include a non-exhaustive search in the search space. Thus, initialization in the search space can be an important part of the encoding process, which provides a starting point that may produce improved calculation results within a given processing time.
The present invention also describes motion estimation techniques that can be used to quickly identify acceptable motion vectors after initialization. This motion estimation technique can perform a non-exhaustive search of the search space to limit the number of calculations performed. In one example, the motion estimation routine may include defining a circle with a radius (R) around the initialized pixel position, and the current video block to be encoded and the video in the search space associated with a set of pixel positions within the circle. Block comparison. The routine may also include identifying the pixel position of the video block that produces the lowest difference in the circle, and when the identified pixel position corresponds to the center of the circle, encoding the current video block using the motion vector defined by the identified pixel position .
If the identified pixel position of the video block that produces the lowest difference in the circle does not correspond to the center of the circle with radius (R), a new circle with radius (R') can be defined around the pixel position . The radius (R) may be equal to the radius (R'), although the present technology is not limited to this relationship. In either case, the motion estimation routine may also include determining the pixel position of the video block that produces the lowest difference within a circle with a radius (R), and when the pixel position corresponds to a radius (R) At the center of the circle, use the motion vector defined by the pixel position to encode the current video block. The motion estimation routine can continue in a similar manner, defining additional circles as needed, until the video block that produces the lowest difference corresponds to the center of the circle. In many cases, the video block that produces the lowest difference can be spatially close to the initial estimate. By defining the circle until the video block that produces the lowest difference corresponds to the center of the circle, the acceptable motion vector can be accurately located very quickly, and unnecessary comparisons and calculations can be reduced.
The present invention also describes other coding techniques that can be used to improve real-time coding, especially when computing resources are limited. For example, during the video encoding process of video frames, the available computing resources can be monitored. If the earlier video blocks of the video frame are encoded very quickly, the saved calculations can be identified. In this case, if desired, more exhaustive search techniques can be used to encode later video blocks of the same video frame, because the saved calculations associated with earlier video blocks can be provided to later video blocks The encoding used.
The techniques described in this invention can provide one or more advantages. For example, these techniques can help reduce power consumption during video encoding, and can facilitate and effective real-time video encoding by accelerating motion estimation routines. These technologies can also improve video coding in terms of video quality and compression, especially when computing resources are limited. Specifically, these techniques can increase the degree of data compression compared to some conventional search techniques such as the conventional diamond search technique. The encoding technology described in the present invention is particularly useful in wireless digital video equipment or other portable video equipment where computing resources and bandwidth are more limited, and power consumption is very important.
FIG. 1 shows a block diagram of an example system 2 in which a source device 4 sends an encoded video data sequence to a sink device 8 through a communication link 6. The source device 4 and the sink device 8 are both digital video devices. Specifically, the source device 4 uses any one of a variety of video compression standards such as MPEG-4 developed by the Moving Picture Experts Group to encode and transmit video data. Other standards may include MPEG-1, MPEG-2 or other MPEG standards developed by the Motion Picture Experts Group, ITU H.263 and similar standards, QuickTimeTM technology developed by Apple Inc. of Cupertino, California, and Video for WindowsTM developed by Microsoft Corporation of Redmond, State, IndeoTM developed by Intel Corporation, and CinepakTM developed by SuperMac Corporation.
The communication link 6 may include a wireless link, a physical transmission line, a partition-based network such as a local area network, a wide area network, or a global network such as the Internet, a public switched telephone network (PSTN), or a combination of various links and networks. In other words, the communication link 6 represents any suitable communication medium for sending video data from the source device 4 to the sink device 8, or may be a combination of different networks and links.
The source device 4 may be any digital video device capable of encoding and transmitting video data. For example, the source device 4 may include a memory 22 for storing a digital video sequence, a video encoder 20 for encoding the sequence, and a transmitter 14 for transmitting the encoded sequence through the communication link 6. For example, a video encoder may include a digital signal processor (DSP) that executes programmable software modules that define encoding techniques.
The source device 4 may also include an image sensor 23 for capturing a video sequence and storing the captured sequence in the memory 22, such as a video camera. In some cases, the source device 4 can send a real-time video sequence through the communication link 6. In those cases, the receiving device 8 may receive the real-time video sequence and display the video sequence to the user.
The receiving device 8 may be any digital video device capable of receiving and decoding video data. For example, the receiving device 8 may include a receiver 15 for receiving an encoded digital video sequence, a decoder 16 for decoding the sequence, and a display screen 18 for displaying the sequence to the user.
Examples of source device 4 and sink device 8 include servers, workstations or other desktop computing devices located on a computer network, and mobile computing devices such as laptop computers or personal digital assistants (PDAs). Other examples include digital television broadcasting satellites and receiving devices such as digital televisions, digital cameras, digital video cameras or other digital recording devices, digital video phones such as cellular wireless phones and satellite wireless phones with video capabilities, other wireless video devices, etc. .
In some cases, each of the source device 4 and the sink device 8 includes an encoder/decoder (codec) (not shown) for encoding and decoding digital video data. In this case, both the source device 4 and the sink device 8 may include a transmitter and a receiver, as well as a memory and a display screen. Many of the encoding techniques described later are described in the context of digital video equipment including encoders. However, it is understood that the encoder may form part of the codec. In this case, the codec can be implemented by a DSP, a microprocessor, an application specific integrated circuit (ASIC), discrete hardware components, or various combinations thereof.
For example, source device 4 may operate on pixel blocks within a sequence of video frames to encode video data. For example, the video encoder 20 of the source device 4 may perform a motion estimation coding technique in which a video frame to be transmitted is divided into pixel blocks (referred to as video blocks). For each video block in the video frame, the video encoder 20 of the source device searches the video block stored in the memory 22 to find the previous video frame (or the next video frame) that has been sent to identify similar videos Block, and encode the difference between the video blocks, and identify the motion vector of the video block of the previous frame (or the next frame) used for encoding. The motion vector may define the pixel position associated with the upper left corner of the video block, although other formats of motion vector may also be used. In either case, by using motion vectors to encode video blocks, the bandwidth required to transmit the video data stream can be significantly reduced. In some cases, the source device 4 can support a programmable threshold, which can prompt the termination of various comparisons or calculations during the encoding process to reduce the number of calculations and save power.
The receiver 15 of the receiving device 8 can receive encoded video data in formats such as motion vector and encoded difference. The decoder 16 performs motion compensation technology to generate a video sequence for display to the user through the display screen 18. The decoder 16 of the receiving device 8 may also be implemented as an encoder/decoder (codec). In this case, the source device 4 and the sink device 8 can encode, send, receive, and decode digital video sequences.
Shown in Figure 2 is a block diagram of an example digital video device 10, such as a source device 4 incorporating a video encoder 20 that encodes a digital video sequence in accordance with one or more of the techniques described in the present invention. The exemplary digital video device 10 shown is a wireless device, such as a mobile computing device, a personal digital assistant (PDA), a wireless communication device, a wireless phone, and so on. However, the technology in the present disclosure is not necessarily limited to wireless devices, and can be easily applied to other digital video devices including non-wireless devices.
In the example of FIG. 2, the digital video device 10 transmits the compressed digital video sequence through the transmitter 14 and the antenna 12. The video encoder 20 encodes the video sequence before sending, and buffers the encoded digital video sequence in the video memory 22. For example, as described above, the video encoder 20 may include a programmable digital signal processor (DSP), a microprocessor, one or more application specific integrated circuits (ASIC), dedicated hardware components, or various combinations of these devices and components. The memory 22 can store computer-readable instructions and data, which are used by the video encoder 20 in the encoding process. For example, the memory 22 may include synchronous dynamic random access memory (SDRAM), flash memory, electrically erasable programmable read-only memory (EPROM), and so on.
In some cases, the digital video device 10 includes an image sensor 23 for capturing video sequences. For example, the image sensor 23 can capture video sequences and store them in the memory 22 before encoding. The image sensor 23 can also be directly coupled to the video encoder 22 to improve video encoding in real time. Encoding techniques that implement the initialization routines and/or motion estimation routines described later can speed up the encoding process, reduce power consumption, improve data compression, and may be beneficial to real-time video encoding of devices with relatively limited processing capabilities.
As an example, the image sensor 23 may include a camera. Specifically, the image sensor 23 may include a charge coupled device (CCD), a charge injection device, a photodiode array, a complementary metal oxide semiconductor (CMOS) device, or any other photosensitive device capable of capturing video images or digital video sequences.
The video encoder 20 may include an encoding controller 24 that controls the execution of encoding algorithms. The video encoder 20 may also include a motion encoder 26 that performs the motion estimation techniques described in the present invention. If necessary, the video encoder 20 may also include additional components such as a texture encoder (not shown) that perform intra-frame compression commonly used to compress still images, such as discrete cosine transform (DCT) encoding. For example, texture coding can be performed in addition to motion estimation, or it may be used instead of motion estimation if the processing power is deemed too limited for effective motion estimation.
The components of the video encoder 20 may include software modules executed on the DSP. Optionally, one or more of the video encoders 20 may include hardware components specifically designed to perform one or more aspects of the techniques described in the present invention, or one or more ASICs.
FIG. 3 shows an example video block in the form of a macro block 31 that can be stored in the memory 22. The MPEG standard and other video coding modes can use video blocks in the macroblock mode in the motion estimation video coding process. In the MPEG standard, the term "macroblock" refers to a 16×16 set of pixel values from a subset of video frames. Each pixel value can be represented by one byte of data, although a larger or smaller number of bits can also be used to define each pixel to achieve the desired imaging quality. The macro block may include several smaller 8×8 pixel image blocks 32. However, generally speaking, the encoding technique described in the present invention can be operated using blocks of any certain size, such as a 16-byte×16-byte macro block, an 8-byte×8-byte image block, or if necessary, Use video blocks of different sizes.
Figure 4 shows an example search space 41 that can be stored in memory. The search space 41 is a collection of video blocks corresponding to a previously transmitted video frame (or the next video frame of a series of video frames). The search space may include the entirety of the previous or intended video frame, or, if necessary, a subset of the video frame. As shown in the figure, the search space can be rectangular, or any of various shapes and sizes can be assumed. In either case, during the video encoding process, the current frame to be encoded is compared with a block in the search space to identify an appropriate match, so that the difference between the current block and a similar video block in the search space can be identified together with the The motion vectors of similar video blocks are sent together.
The initialization technique can take advantage of the phenomenon of spatial redundancy, so that the initial comparison of the current video block 42 and the video block in the search space 41 is more likely to identify an acceptable match. For example, the technique described later can be used to identify the pixel location 43, which can identify the video block in the search space that is initially compared with the video block 42. Initializing the position close to the best matching video block can increase the likelihood of finding the best motion vector and can reduce the average number of searches required to find an acceptable motion vector. For example, initialization can provide an improved starting point for the motion estimation routine, as well as improved calculation results within a given processing time period.
In the motion estimation video encoding process, the motion encoder 26 may use a comparison technique such as a sum of absolute difference (SAD) technique or a sum of square difference (SSD) technique to compare the current frame to be encoded with the previous video block. Other comparison techniques can also be used.
The SAD technology involves performing the task of comparing the absolute difference between the pixel value of the current block to be encoded and the pixel value of the previous block of the current block to which it is compared. The results of these absolute difference comparisons are added, that is, accumulated, to define a difference indicating the difference between the current video block and the previous video block to which the current video block is compared. For an 8×8 pixel image block, 64 differences can be calculated and added, and for a 16×16 pixel macro block, 256 differences can be calculated and added. A lower difference value generally indicates that the video block compared with the current video block is a better match, and thus is a better candidate block that can be used in the motion estimation coding process than the video block that produces the higher difference value. In some cases, when the accumulated difference exceeds a defined threshold, the calculation can be terminated. In this case, no additional calculations are needed, because the video block compared with the current video block cannot be accepted for efficient motion estimation encoding. SSD technology also involves performing the pixel value of the current block to be encoded and comparing it with the current block. The task of calculating the difference between the pixel values of the previous block of the block. In SSD technology, the result of the absolute difference calculation is squared, and then the squared values are added, that is, accumulated, to define a difference indicating the difference between the current video block and the previous video block compared with the current block. Optionally, other comparison techniques such as mean square error (MSE), standardized cross-correlation function (NCCF), or another suitable comparison algorithm can be implemented.
Figure 5 shows a flowchart of a video coding technique that can implement an initialization motion estimation routine in a digital video device. In the example of FIG. 5, the video encoder 20 uses the median function to calculate the initialized pixel position in the search space. However, other linear or non-linear functions can be used to apply the same principle, such as averaging or weighting functions.
As shown in FIG. 5, in order to initialize the encoding of the current video block, the encoding controller 24 calculates the median X coordinate pixel position of the motion vector of the adjacent video block (51). Similarly, the encoding controller 24 calculates the median Y coordinate pixel position of the motion vector of the adjacent video block (52). In other words, the previously calculated motion vectors associated with a set of video blocks located at a certain position relative to the current video block to be encoded are used to calculate the median X and Y coordinates. Adjacent video blocks may form a group of video blocks located at a certain position relative to the current video block. For example, depending on the implementation, the set of video blocks can be defined to include or exclude certain video blocks. In either case, the motion estimation of the current current video block is then initialized on the median X and Y coordinates (53). The motion estimator 26 then uses conventional search techniques or search techniques similar to those described in detail below to search for acceptable motion vectors for the current video block (54).
Fig. 6 is a conceptual illustration of an exemplary video frame in which a current video block is encoded using a technique similar to that shown in Fig. 5. Specifically, a previously calculated motion vector associated with a group of video blocks located at a certain position relative to the current video block 63 may be used to initialize the motion estimation of the current video block 63. In FIG. 6, the group of video blocks located at a certain position includes neighboring block 1 (62A), neighboring block 2 (62B), and neighboring block 3 (62C). However, as mentioned, the group can be defined in various other formats, including or excluding video blocks located at various positions relative to the current video block 63.
In the example of FIG. 6, the initialization position for performing motion estimation of the current block 63 in the search space can be defined by the following formula: X initial = median (XMV1, XMV2, XMV3) Y initial = median (YMV1, YMV2 , YMV3) Optionally, an average value function can be used. In this case, the search space can be defined by the following formula to perform the initialization position of the motion estimation of the current block 63: X initial = average value (XMV1, XMV2, XMV3 )Yinitial=average value (YMV1, YMV2, YMV3) Other linear or non-linear mathematical formulas or relationships between adjacent video blocks and the current video block to be encoded can also be used.
In either case, by defining the location of the initialization of the motion estimation search, encoding can be accelerated by increasing the probability of quickly identifying an acceptable match between the video block to be encoded and the video block in the search space. As shown in FIG. 7 (compared to FIG. 6), adjacent video blocks forming a group of video blocks for generating the initialization position include blocks in other determined positions relative to the position of the current video block 63. In many cases, the position of the neighboring block 62 relative to the current block 63 may depend on the direction of the video block of the encoded frame 61.
For example, the coded neighboring video block of the current video block 63 generally has a calculated motion vector, while other neighboring video blocks that have not been coded generally do not have a motion vector. Thus, if the encoding of the video block proceeds from left to right and top to bottom, starting from the upper leftmost video block of frame 61, the neighboring block 62A including a subset of the neighboring video blocks of the current video block 63 can be used. -62C (Figure 6). Optionally, if the encoding of the video block is performed from right to left and bottom to top, starting from the bottom right video block of frame 61, a different subset of adjacent video blocks including the current video block 63 may be used The adjacent blocks 62D-62F (Figure 7). According to the principles of the present disclosure, many other variations of the video block group used to generate the initialization position can be defined. For example, if a very high level of motion occurs in the video sequence, a determined video block position that is farther from the video block 63 than the immediately adjacent block may be included in the group.
The initialization technique of FIG. 5 may not be used for the first video block of the frame 61 to be encoded, because at this point, the motion vectors of adjacent video blocks are not available. However, once one or more motion vectors of neighboring video blocks are identified, the initialization routine can be started at any time thereafter. In a specific case, the initialization technique can start the encoding of the second video block in the second row. Before this point, the motion vectors of the three adjacent blocks are usually not defined. However, in other cases, the initialization routine may start at any time when the motion vector is calculated for at least one neighboring block of the video block to be encoded.
As shown in FIG. 8, the motion vectors associated with a larger number of neighboring blocks of the current video block 63 may be used to calculate the initialization position in the search space. Since more motion vectors are calculated for more neighboring blocks, more motion vectors can be used to calculate the initial position for performing motion estimation on the current video block. In some cases, in order to more fully utilize the spatial correlation phenomenon, when calculating the initial position, the motion vector associated with the neighboring block spatially adjacent to the current video block 42 may be given to the motion vector that is spatially distant from the current video block 42. The motion vector associated with the neighboring block has a larger weight. In other words, a weighting function such as a weighted average value or a weighted median value can be used to calculate the initial pixel position. For example, the motion vectors associated with the neighboring blocks 1-4 (62A, 62B, 62C, and 62G) can be given greater weights than the motion vectors associated with the neighboring blocks 5-10 (62H-62M) because the neighbors Blocks 1-4 are statistically more likely to be video blocks used in motion estimation coding. These and other possible modifications can be made clear by reading this disclosure.
Figure 9 shows a flowchart of a video coding technique that can be implemented in a digital video device to perform motion estimation. This technique may involve non-exhaustive search of video blocks in the search space to reduce the number of calculations required for video encoding. As shown in the figure, the encoding controller 24 initializes the motion estimation based on the motion vector calculated for the neighboring video blocks of the frame (91). For example, the initialization process may include a process similar to that shown in FIG. 5, wherein the adjacent video block used for initialization includes a group of video blocks located at a certain position relative to the current video block to be encoded.
Once the initialization position is calculated, the motion estimator 26 identifies a set of video blocks defined by the pixel positions within the circle of the radius (R) of the initialization position (92). For example, the radius (R) may be large enough to define a circle including at least five pixels, although a larger radius may also be defined. More preferably, the radius (R) is large enough to define a circle including at least 9 pixels, that is, the initialization position and all eight pixel positions closely surrounding the initialization position. The inclusion of the initialization position and all eight pixel positions closely surrounding the initialization position can improve the search technique by expecting a motion vector in every possible direction relative to the initialization position.
The motion estimator 26 then compares the current video block to be encoded with the video block defined by the pixel positions within the circle (93). The motion estimator 26 may then identify the video block that produces the lowest difference within the circle. If the video block defined by the center of the circle with the right radius (R) produces the lowest difference, that is, the lowest difference metric as defined by the comparison technique used ("Yes" branch of 94), then the motion estimator 26 can use Identify the motion vector of the video block defined by the center of the circle with radius (R) to encode the current video block (95). For example, the SAD or SSD comparison described above can be used. However, if the video block defined by the center of the circle with the right radius (R) does not produce the lowest difference ("No" branch of 94), the motion estimator 26 identifies the pixel position of the video block that produces the lowest difference by identifying The video block (96) defined by the pixel location within the radius (R').
The motion estimator 26 then compares the video block to be encoded with the video block defined by the pixel positions within the circle of radius (R'). The motion estimator 26 may identify the video block that produces the lowest difference within a circle with a radius (R). If the video block defined by the center of the circle with radius (R) produces the lowest difference, that is, the lowest difference metric as defined by the comparison technique used ("Yes" branch of 94), the motion estimator 26 uses the flag The motion vector of the video block defined by the center of the circle with radius (R') encodes the current video block (95). However, if the video block defined by the center of a circle with a radius of (R') does not produce the lowest difference ("No" branch of 94), the motion estimator 26 continues to define another circle with a radius of (R"). The process, and so on.
If necessary, each subsequent circle defined around the pixel location that produces the lowest difference may have the same radius as the previous circle, or a different radius. By defining the radius of the most recently identified pixel position relative to the video block that produces the lowest difference in the identification search space, the best video block used in motion estimation can be quickly identified without the need for exhaustive search in the search space. . In addition, the simulation shows that the search technique shown in FIG. 9 can achieve improved compression relative to the conventional diamond search technique, which operates to associate the difference of the pixel block with the pixel located in the center of a diamond group of pixels. minimize.
10-13 are conceptual illustrations of a search space in which a video coding technique similar to that shown in FIG. 9 is performed to find acceptable motion vectors. Each point in the grid represents a pixel location that identifies a unique video block in the search space. The motion estimation routine can be initialized at the pixel position (X5, Y6), such as by executing the initialization routine shown in FIG. 5. After finding the initial position (X5, Y6), the motion estimator 26 defines a circle with a radius (R), and compares the video block in the search space associated with the pixel position in the circle with the radius (R) with the video block in the search space. Compare with the current video block of the encoding. The motion estimator 26 then identifies the video block that produces the lowest difference within a circle of radius (R).
If the motion estimator 26 determines that the pixel location (X6, Y7) identifies the video block that produces the lowest difference, the motion estimator 26 defines a circle with a radius (R') around the pixel location (X6, Y7), as shown in Figure 11 Show. The motion estimator 26 then compares the video block in the search space associated with the pixel position within the circle of radius (R) with the current video block to be encoded. If the motion estimator 26 determines that the pixel location (X6, Y8) identifies the video block that produces the lowest difference, the motion estimator 26 defines another circle with a radius (R") around the pixel location (X6, Y8), as shown in the figure 12. The motion estimator 26 then compares the video block in the search space associated with the pixel position within the circle of radius (R") with the current video block to be encoded. If the motion estimator 26 determines that the pixel position (X7, Y9) identifies the video block that produces the lowest difference, the motion estimator 26 defines another circle with a radius of (R) around the pixel position (X7, Y9). The motion estimator 26 then compares the video block in the search space associated with the pixel position within the circle of radius (R') with the current block to be encoded. Finally, the motion estimator 26 should locate the pixel location corresponding to the center of the circle that produces the lowest difference. At this point, the motion estimator can use the pixel position that also corresponds to the center of the circle that produces the lowest difference as the motion vector of the current video block.
When defining each new circle, only comparisons associated with pixel positions not included in the previous circle need to be performed. In other words, referring to Figure 11, for a circle with a radius of (R'), at this point, only the pixel positions (Y8, X5), (Y8, X6), (Y8, X7), (Y7, X7) The video block associated with (Y6, X7). The video blocks associated with other pixel positions in the circle of radius (R), namely (Y7, X5), (Y7, X6), (Y6, X5) and (Y6, X6), have been associated with the radius of (R ) Of the circle (Figure 10) is executed.
The radii (R), (R'), (R"), (R), etc. can be equal to or unequal to each other, depending on the implementation. Also, in some cases, the radius can be defined so that each group includes A larger number of pixels, which can improve the likelihood of finding the best motion vector. However, a larger radius will increase the number of calculations for any given search. However, in either case, the Defining a circle around each pixel can have advantages over other techniques. For example, when defining a circle around the pixel location, each neighboring pixel location can be checked during the comparison. In other words, all eight pixels that closely surround the center pixel The position can be included in the defined circle. The exhaustive search that requires comparison of every possible video block defined by every pixel in the search space can still be avoided, which can obtain accelerated video coding.
Figure 14 shows a flowchart of another video coding technique that can be implemented in digital video equipment to improve real-time video coding. As shown in the figure, the encoding controller 24 initializes the motion estimation of the current video block of the video frame to be encoded (141). For example, the initialization process may include a process similar to that shown and described with reference to FIGS. 5-8, in which the initial pixel position in the search space is identified. In addition, the initialization process may include defining the scope of the search that may be performed to encode the current video block. If the current video block is the first video block to be encoded for a video frame, the search range can be defined by a default value programmed into the encoding controller 24.
After initialization, the motion estimator 26 searches for a motion vector in the search space to encode the current video block (142). For example, the search and encoding process may include a process similar to that shown and described with reference to FIGS. 9-13, in which a circle is defined around the pixel position until the pixel position that produces the lowest difference corresponds to the center of the circle. The search range can be limited, such as by limiting the amount of time that can be searched, or by limiting the size of the circle used. Again, the search range can be defined through the above initialization.
If the video frame to be encoded includes other video blocks ("Yes" branch of 143), the encoding controller 24 identifies the remaining amount of computing resources for the frame (144). The encoding controller 24 then initializes the motion estimation of the subsequent video block of the video frame to be encoded (141). If computational savings are achieved in the encoding of an earlier video block, the amount of available resources can be increased compared to the resources available for encoding the previous video block. The computing resources can be defined by the clock speed of the video encoder, and the desired resolution of the video sequence is sent according to the number of frames per second.
When the encoding controller 24 initializes the motion estimation of the subsequent video block of the video frame to be encoded, the initialization process again includes the process of defining the search range. However, in this case, the search range may be based on the identified amount of available computing resources. Thus, if the encoding of the first video block is performed very quickly, a more exhaustive search can be used to encode subsequent video blocks. For example, when more computing resources are available, the radius used to define a circle around the pixel location can be increased. These or similar technologies can improve the quality of real-time video encoding when computing resources are limited, for example, by improving the compression ratio of encoded video blocks. Once all the video blocks of the video frame are encoded (the "No" branch of 143), the device 10 may transmit the encoded frame through the transmitter 14 and the antenna 12 (145).
Table 1 lists the data collected during the simulation of the technology described above relative to other conventional technologies. Coding is implemented on a video sequence with a relatively large number of motions. A conventional diamond search technique and an exhaustive search technique in which all video blocks in the search space are compared with the video block to be coded are also used to encode the same video sequence. For each technique, Table 1 lists the file size, the signal-to-noise ratio, the average number of searches per macro block, the maximum number of searches per macro block, and the number of searches required for the frame in the worst case. The label "circle search" refers to an encoding technique that uses a process similar to that of Figures 5 and 9. As can be understood from Table 1, the technology described in the present invention can achieve improved compression with a reduced number of searches compared to the conventional diamond search. The technology described in the present invention may not be able to achieve the compression level of the full search. However, in this example, the circle search technique only requires an average of 17.2 searches per macroblock, which is in contrast to the diamond search requiring 21.3 searches per macroblock and the full search requiring 1024 searches per macroblock.
Table 1
Numerous different embodiments are described. For example, a video coding technique for initializing a search in a search space, and a motion estimation technique for performing the search are described. These technologies can improve video encoding by avoiding calculations in certain situations, speeding up the encoding process, improving compression, and possibly reducing power consumption in the video encoding process. In this way, these technologies can improve video coding in accordance with standards such as MPEG-4, and can better facilitate the implementation of video coding in wireless devices where computing resources are limited and power consumption is very important. In addition, these technologies will not affect interoperability with decoding standards such as the MPEG-4 decoding standard.
However, various modifications can be made without departing from the scope of the appended claims. For example, the initialization routine can be extended to calculate multiple initial pixel positions within the search space. For example, different linear or non-linear functions may be used to calculate two or more initialization positions based on the number of motions of a group of video blocks relative to the determined position of the video block to be encoded. Likewise, two or more different video block groups located at a certain position relative to the video block to be encoded may be used to calculate two or more initialization positions. In some cases, the calculation of two or more initialization positions can further speed up the encoding process. These and other modifications can be made clear by reading this disclosure.
The technology described in the present invention can be implemented in hardware, software, firmware or any combination thereof. If implemented in software, these technologies can target computer-readable media including program codes. When executed in a device that encodes a video sequence conforming to the MPEG-4 standard, one or more of the methods described above will be executed. A. In this case, the computer-readable medium may include random access memory (RAM), such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), and Erase programmable read-only memory (EEPROM), flash memory, etc.
The program code can be stored in the memory in the form of computer readable instructions. In this case, a processor such as a DSP can execute instructions stored in a memory to implement one or more of the techniques described in the present invention. In some cases, these techniques can be executed by a DSP that calls various hardware components to speed up the encoding process. In other cases, the video encoder can be implemented by a microprocessor, one or more application specific integrated circuits, one or more field programmable gate arrays (FPGA), or some other hardware-software combination. These and other embodiments are included within the scope of the appended claims.
14 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9307239B2 | Cited by | United States of America | Applicant |
| US9525879B2 | Cited by | United States of America | Applicant |
| US7864837B2 | Cited by | United States of America | Applicant |
| US9300963B2 | Cited by | United States of America | Applicant |
| CN103299630A | Cited by | China | Search report |
| CN111818333A | Cited by | China | Search report |
| US9860552B2 | Cited by | United States of America | Applicant |
| WO2012122927A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US8755437B2 | Cited by | United States of America | Applicant |
13 members in 9 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 10176028 | United States of America | – | |
| 17602802 | United States of America | A | |
| 17602802 | United States of America | A | |
| 10176028 | – | – | – |
| US20020176028 | – | – | – |
Members13
| Document | Office | Kind | |
|---|---|---|---|
| US2003231712A1 | United States of America | A1 | |
| WO03107680A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU2003238294A1 | Australia | A1 | |
| TW200402236A | Taiwan Province of China | A | |
| WO03107680A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20050012794A | Republic of Korea | A | |
| EP1514424A2 | European Patent Office (EPO) | A2 | |
| CN1663280AThis record | China | A | |
| JP2005530421A | Japan | A | |
| TWI286906B | Taiwan Province of China | B | |
| MY139912A | Malaysia | A | |
| KR100960847B1 | Republic of Korea | B1 | |
| US7817717B2 | United States of America | B2 |
5 legal events, as 2 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Rejection of a patent application after its publicationC12 | C12 | CN | |
| Applications withdrawn, deemed to be withdrawn, or refused after publication in hong kongWithdrawnWD | WD | HK | |
| Requests to designate patent in hong kongDE | DE | HK | |
| Entry into substantive examinationC10 | C10 | CN | |
| PublicationC06 | C06 | CN |
Numbers
- Publication
- 1663280
- Publication, DOCDB
- 1663280
- Publication, EPODOC
- CN1663280
- Application
- 38142783
- Application, DOCDB
- 03814278
- Application, EPODOC
- CN2003814278
Titles2
- Chinese
- 用于视频编码的运动估计技术
- English
- Motion estimation technology for video coding
Classification
- CPC, 5
- H04N19/51
- H04N19/127
- H04N19/56
- H04N19/557
- H04N19/57
- IPC, 3
- G06T9 00
- H03M7 36
- H04N19 51