EP3470533B1

Systems and methods to detect copy number variation

Abstract

This record has no abstract on file.

EP3470533B1, drawing sheet 1
Sheet 1 of 16

Term

6.9 yearsleft in the term

Expires 4 September 2033.

  1. Priority and filed
  2. Granted
  3. Today
  4. Expires

13 claims: 9 independent, 4 dependent

  1. 1
    A method of detennining copy number variation in a sample that includes cell-free polynucleotides, the method comprising:a. providing at least two sets of cell-free polynucleotides, which map to different mappable positions in a reference sequence in a genome, and, for the sets of cell-free polynucleotides;i. non-uniquely tagging the cell-free polynucleotides with a set of molecular barcodes;ii. amplifying the cell-free polynucleotides to produce amplified polynucleotides;iii. sequencing a subset of the set of amplified polynucleotides, to produce a set of sequencing reads;iv. grouping the set of sequencing reads sequenced from amplified polynucleotides into families which correspond to sequencing reads of polynucleotides amplified from the same cell-free polynucleotide;v. inferring a quantitative measure of families in the sets;and b. detennining copy number variation based on the quantitative measure of families in the sets.
  2. 4
    The method of any one of claims 1-3, wherein the molecular barcodes are attached to the cell-free polynucleotides through an enzymatic reaction such as a ligation reaction.
  3. 6
    The method of any one of claims 1-5, further comprising selectively enriching regions from a genome or transcriptome of the subject prior to sequencing.
  4. 7
    The method of any one of claims 1-6, further comprising filtering out sequencing reads with an accuracy or quality score of less than a threshold and/or mapping score of less than a threshold.
  5. 8
    The method of any one of claims 1-7, wherein the sequencing reads are grouped into families based on the non-unique barcode sequence in combination with the sequence data at the beginning (start) and end (stop) portions of the sequencing reads, optionally further combining the length of the sequencing reads.
  6. 9
    The method of any one of claims 1-8, wherein inferring a quantitative measure of families in the set comprises determining the number of families mapping to different reference loci.
  7. 11
    The method of any one of claims 1-9, wherein the quantitative measure is normalized for representational bias during the sequencing process.
  8. 12
    The method of any one of claims 1-11, wherein the quantitative measure is a count.
  9. 13
    A computer readable medium comprising non-transitory machine-executable code that, upon execution by a computer processor, implements a method, the method comprising:a. accessing a data file comprising a plurality of sequencing reads, wherein the sequence reads derive from progeny polynucleotides amplified from non-uniquely tagged parent cell-free polynucleotides;b. grouping sequencing reads sequenced from the progeny polynucleotides into families comprising sequencing reads of progeny polynucleotides amplified from the same tagged parent cell-free polynucleotide;c. inferring a quantitative measure of families in the non-uniquely tagged parent cell-free polynucleotides;and d. determining copy number variation by comparing the quantitative measure of families in the non-uniquely tagged parent cell-free polynucleotides.