EP1501047A2

Fingerprinting segments of data content for version identification

Abstract

A method of detecting a version of input data content in which the data content is arranged as two or more segments according to a segmentation pattern; and versions of the data content are identifiable by corresponding identification data patterns by which at least some of the segments have respective identification data; comprises the steps of detecting identification data in respect of each segment of the input data content; comparing the detected identification data with identification data patterns corresponding to different versions of the data content; and detecting that the input data content comprises at least a contribution from a certain version of the data content if a sum of matches obtained between the detected identification data and the identification data pattern for that version exceeds a threshold number.

EP1501047A2, drawing sheet 1
Sheet 1 of 19

Term

Term ended

Projected expiry passed 16 July 2024, 2.2 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

31 claims: 14 independent, 17 dependent

  1. 1
    A method of detecting a version of input data content in which:the data content is arranged as two or more segments according to a segmentation pattern;and versions of the data content are identifiable by corresponding identification data patterns by which at least some of the segments have respective identification data;the method comprising the steps of: detecting identification data in respect of segments of the input data content;comparing the detected identification data with identification data patterns corresponding to different versions of the data content;and detecting that the input data content comprises at least a contribution from a certain version of the data content if a sum of matches obtained between the detected identification data and the identification data pattern for that version exceeds a threshold number.
  2. 5
    A method according to any one of the preceding claims, comprising the step of:weighting a match between identification data detected in respect of a segment of the input data content according to the number of instances of identification data detected in respect of that segment of the input data content, the sum of matches being a weighted sum of matches.
  3. 8
    A method according to any one of the preceding claims, comprising the step of:if identification data is not detected in respect of two or more segments of the input data content, combining those segments into groups of two or more segments and detecting identification data in respect of the combined groups of segments.
  4. 10
    A method according to any one of the preceding claims, in which the threshold number represents a number of segments less than the total number of segments.
  5. 11
    A method according claim 10, in which the threshold number represents a number of segments less than the total number of segments having associated identification data in that identification data pattern.
  6. 12
    A method according to any one of the preceding claims, in which versions of the data content are identifiable by corresponding identification patterns by which substantially all of the segments have respective identification data.
  7. 13
    A method of applying identification data to input data content, the method comprising the steps of:(i) generating n instances of the input data content, where n is greater than one, at least all but one of the instances carrying respective identification, the identification data of each of the instances carrying respective identification data being unique with respect to the respective identification data carried by the other instances;and (ii) generating versions of the input data content by selecting segments from the n instances, so that each version of the input data content carries identification data from the instances in accordance with an associated identification data pattern;followed by one or more iterations of the steps of: (iii) generating m further instances of the input data content, where m is one or more, each of the m instances carrying respective identification data which is unique with respect to all of the other instances;and (iv) generating further versions of the input data content by selecting segments from the m instances, a set of instances including the m instances, or the entire set of generated instances, so that each version of the input data content carries identification data from the instances in accordance with an associated identification data pattern.
  8. 18
    A method of applying identification data to input data content, the method comprising the steps of:(i) providing n instances of the input data content, where n is greater than one, at least all but one of the instances carrying respective identification data, the identification data of each of the instances carrying respective identification data being unique with respect to the respective identification data carried by the other instances;and (ii) generating versions of the input data content by selecting segments by a predetermined segmentation pattern from the n instances, so that each version of the input data content carries identification data from the instances in accordance with an associated identification data pattern;in which the segmentation pattern is such that at least one of the segments is not contiguous within the input data content.
  9. 20
    A method according to any one of the preceding claims, in which the data content comprises video content having a plurality of successive images.
  10. 23
    Computer software having program code for carrying out a method according to any one of the preceding claims.
  11. 27
    Apparatus for detecting a version of input data content in which:the data content is arranged as two or more segments according to a segmentation pattern;and versions of the data content are identifiable by corresponding identification data patterns by which at least some of the segments have respective identification data;the apparatus comprising: means for detecting identification data in respect of segments of the input data content;means for comparing the detected identification data with identification data patterns corresponding to different versions of the data content;and means for detecting that the input data content comprises at least a contribution from a certain version of the data content if a sum of matches obtained between the detected identification data and the identification data pattern for that version exceeds a threshold number.
  12. 28
    Apparatus for applying identification data to input data content, the apparatus comprising:(i) instance generation means for generating n instances of the input data content, where n is greater than one, at least all but one of the instances carrying respective identification data, the identification data of each of the instances carrying respective identification data being unique with respect to the respective identification data carried by the other instances;and (ii) version generation means for generating versions of the input data content by selecting segments from the n instances, so that each version of the input data content carries identification data from the instances in accordance with an associated identification data pattern;(iii) means for controlling the instance generation means to generate m further instances of the input data content, where m is one or more, each of the m further instances carrying respective identification data which is unique with respect to all of the other instances;and (iv) means for controlling the version generation means to generate further versions of the input data content by selecting segments from the m instances, a set of instances including the m instances, or the entire set of generated instances, so that each version of the input data content carries identification data from the instances in accordance with an associated identification data pattern.
  13. 29
    Apparatus for applying identification data to input data content, the apparatus comprising:(i) means for providing n instances of the input data content, where n is greater than one, at least all but one of the instances carrying respective identification data, the identification data of each of the instances carrying respective identification data being unique with respect to the respective identification data carried by the other instances;and (ii) means for generating versions of the input data content by selecting segments by a predetermined segmentation pattern from the n instances, so that each version of the input data content carries identification data from the instances in accordance with an associated identification data pattern;in which the segmentation pattern is such that at least one of the segments is not contiguous within the input data content.
  14. 30
    A storage medium carrying data content having associated identification data, the data content comprising segments according to a predetermined segmentation pattern, the segments carrying respective identification data in accordance with an associated identification data pattern, in which the segmentation pattern is such that at least one of the segments is not contiguous within the input data content.