IL227674A

Encoding of video stream based on scene type

Abstract

This record has no abstract on file.

Term

No projected expiry on record.

  1. Priority
  2. Filed
  3. Published
  4. Today

30 claims: 4 independent, 26 dependent

  1. 1
    227674/2 20 CLAIMS:1. A method for encoding a video stream using scene types each having a predefined set of one or more of a plurality of encoder parameters used by a video encoder to encode any given scene type, the method comprising: 5 receiving an input video stream;dividing the input video stream into a plurality of scenes based on scene boundaries, each scene comprising a plurality of temporally contiguous image frames, wherein the dividing comprises determining a given scene boundary according to relatedness of two temporally contiguous image frames in the input video stream, and 10 wherein the determining comprises: scaling one or more high frequency elements of each image frame;converting pixel data in the image frames into frequency coefficients;removing the one or more high frequency elements of each image frame based on the converted frequency coefficients;15 analyzing the image frames to determine a difference between temporally contiguous image frames, wherein a score is computed based on the difference;and identifying a degree of unrelatedness between the image frames when the score exceeds a preset limit, wherein the preset limit score is at a threshold 20 where a scene change occurs;determining scene type for each of the plurality of scenes;and encoding each of the plurality of scenes according to the scene type.
  2. 10
    11. The method for encoding a video stream as recited in claim 10, wherein the performing of the sequential decision-making waterfall process comprises:determining a position of a given scene on a timeline of the input video stream 5 to assign a score based on a predetermined scale according to the position;determining a play-time length of the given scene to assign a score based on a predetermined scale according to the play-time length;determining a motion estimation in the given scene to assign a score based on a predetermined scale according to a magnitude of a motion vector, wherein the motion 10 estimation is a measure of the magnitude of the motion vector;determining a difference in the given scene from a previous scene to assign a score based on a predetermined scale according to the difference;determining a spectral data size of the given scene to assign a score based on a predetermined scale according to the spectral data size;15 identifying facial structures utilizing facial recognition to assign a score based on a predetermined scale according to a number of the facial structures;identifying textual information using optical character recognition in the given scene to assign a score based on a predetermined scale according to an amount of content of the textual information;and 20 determining a level of audience interest from screenplay structure information of the given scene to assign a score based on a predetermined scale according to the level of audience interest. 25 02241443\80-02 227674/2 24
  3. 11
    12. The method of claim 11, wherein the screenplay structure information includes a relative attention parameter, wherein the relative attention parameter approximates a predetermined estimation of a relative amount of viewer attention to be expected for a segment of the input video stream that comprises the given scene.
  4. 16
    17. A video encoding apparatus for encoding a video stream using scene 5 types each having a predefined set of one or more of a plurality of encoder parameters used by the video encoder to encode any given scene type, the apparatus comprising:an input module for receiving an input video stream;a video processing module to divide the video stream into a plurality of scenes based on scene boundaries, each scene comprising a plurality of temporally contiguous 10 image frames, wherein the video processing module determines a given scene boundary according to the relatedness of two temporally contiguous image frames in the input stream, wherein the determining comprises: scaling one or more high frequency elements of each image frame, wherein a transform coder converts pixel data in the image frames into 15 frequency coefficients;removing the one or more high frequency elements of each image frame based on the converted frequency coefficients;analyzing the image frames to determine a difference between temporally contiguous image frames, wherein a score is computed based on the 20 difference;and identifying a degree of unrelatedness between the image frames when the score exceeds a preset limit, wherein the preset limit score is at a threshold where a scene change occurs;the video processing module to determine a scene type for each of the plurality 25 of scenes;and a video encoding module to encode each of the plurality of scenes according to the scene type. 02241443\80-02 227674/2 26
  5. 17
    18. The video encoding apparatus as recited in claim 17, wherein the video processing module determines each scene type based on one or more criteria, the one or more criteria including:a given scene's position on the input video stream's timeline;5 a length of the given scene;a motion estimation in the given scene;a effective difference in the given scene from a previous scene;a spectral data size of the given scene;a optical character recognition in the given scene;or 10 a screenplay structure information of the given scene.
  6. 18
    19. The video encoding apparatus as recited in claim 18, wherein the screenplay structure information utilized by the video encoding apparatus includes a relative attention parameter, wherein the relative attention parameter approximates a predetermined estimation of a relative amount of viewer attention to be expected for a 15 segment of the input video stream that comprises the given scene.
  7. 24
    26. A video encoding apparatus for encoding a video stream using scene 25 types each having a predefined set of one or more of a plurality of encoder parameters used by the video encoder to encode any given scene type, the apparatus comprising:receiving means for receiving an input video stream;02241443\80-02 227674/2 28 dividing means for dividing the input video stream into a plurality of scenes based on scene boundaries, each scene comprising a plurality of temporally contiguous image frames, wherein the dividing means determines a given scene boundary according to the relatedness of two temporally contiguous image frames in the input 5 video stream;determining means for determining a scene type for each of the plurality of scenes based on an assessment in a predetermined scale computed by performing a sequential decision-making waterfall process, wherein the performing of the sequential decision-making waterfall process comprises: 10 determining a position of a given scene on a timeline of the input video stream to assign a score based on a predetermined scale according to the position;determining a play-time length of the given scene to assign a score based on a predetermined scale according to the play-time length;15 determining a motion estimation in the given scene to assign a score based on a predetermined scale according to a magnitude of a motion vector, wherein the motion estimation is a measure of the magnitude of the motion vector;determining a difference in the given scene from a previous scene to 20 assign a score based on a predetermined scale according to the difference;determining a spectral data size of the given scene to assign a score based on a predetermined scale according to the spectral data size;identifying facial structures utilizing facial recognition to assign a score based on a predetermined scale according to a number of the facial 25 structures;identifying textual information using optical character recognition in the given scene to assign a score based on a predetermined scale according to an amount of content of the textual information;and 02241443\80-02 227674/2 29 determining a level of audience interest from screenplay structure information of the given scene to assign a score based on a predetermined scale according to the level of audience interest;and encoding means for encoding each of the plurality of scenes based on the given 5 scene's previously determined encoder parameters that were determined according to the scene type associated with each of the plurality of scenes.
  8. 25
    27. A method for encoding a video stream using scene types each having a predefined set of one or more of a plurality of encoder parameters used by a video encoder to encode any given scene type, the method comprising:10 receiving an input video stream;dividing the input video stream into a plurality of scenes based on scene boundaries, each scene comprising a plurality of temporally contiguous image frames, wherein a given scene boundary is determined according to a screenplay structure information of the input video stream, wherein the screenplay information includes 15 information of story line organization of the input video stream;determining scene type for each of the plurality of scenes based on an assessment in a predetermined scale computed by performing a sequential decisionmaking waterfall process;and encoding each of the plurality of scenes according to the scene type.
  9. 27
    29. The method of claim 27, wherein the screenplay structure information includes a relative attention parameter, wherein the relative attention parameter approximates a predetermined estimation of a relative amount of viewer attention to be expected for each of a plurality of video segments of the input video stream, wherein each of the plurality of video segments could comprise a plurality of scenes. 02241443\80-02 227674/2 30
  10. 30
    32. The method for encoding a video stream as recited in Claim 27, wherein the performing of the sequential decision-making waterfall process comprises:determining a position of a given scene on a timeline of the input video stream 20 to assign a score based on a predetermined scale according to the position;determining a play-time length of the given scene to assign a score based on a predetermined scale according to the play-time length;determining a motion estimation in the given scene to assign a score based on a predetermined scale according to a magnitude of a motion vector, wherein the motion 25 estimation is a measure of the magnitude of the motion vector;determining a difference in the given scene from a previous scene to assign a score based on a predetermined scale according to the difference;determining a spectral data size of the given scene to assign a score based on a predetermined scale according to the spectral data size;02241443\80-02 227674/2 31 identifying facial structures utilizing facial recognition to assign a score based on a predetermined scale according to a number of the facial structures;identifying textual information using optical character recognition in the given scene to assign a score based on a predetermined scale according to an amount of 5 content of the textual information;and determining a level of audience interest from screenplay structure information of the given scene to assign a score based on a predetermined scale according to the level of audience interest. 02241443\80-02