CA2934822C

Methods and systems for detecting genetic variants

Abstract

Disclosed herein in are methods and systems for determining genetic variants (e.g., copy number variation) in a polynucleotide sample. A method for determining copy number variations includes tagging double-stranded polynucleotides with duplex tags, sequencing polynucleotides from the sample and estimating total number of polynucleotides mapping to selected genetic loci. The estimate of total number of polynucleotides can involve estimating the number of double-stranded polynucleotides in the original sample for which no sequence reads are generated. This number can be generated using the number of polynucleotides for which reads for both complementary strands are detected and reads for which only one of the two complementary strands is detected.

CA2934822C, drawing sheet 1
Sheet 1 of 12

Term

8.3 yearsleft in the term

Expires 24 December 2034.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Expires

24 claims: 9 independent, 15 dependent

  1. 1
    CLAIMS WHAT IS CLAIMED IS:1. A method, comprising: (a) providing a sample comprising a set of double-stranded polynucleotide molecules, each double-stranded polynucleotide molecule including first and second complementary strands;(b) tagging said double-stranded polynucleotide molecules with a set of duplex tags, wherein each duplex tag differently tags said first and second complementary strands of a double-stranded polynucleotide molecule in said set;(c) sequencing at least some of said tagged strands to produce a set of sequence reads;(d) reducing redundancy, tracking redundancy, or both in said set of sequence reads;(e) sorting sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to sequence reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in said set, and (ii) each unpaired read represents a first tagged strand having no second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads;(f) determining quantitative measures of (i) said paired reads and (ii) said unpaired reads that map to each of one or more genetic loci;and (g) estimating with a programmed computer processor a quantitative measure of total double-stranded polynucleotide molecules in said set that map to each of said one or more genetic loci based on said quantitative measure of paired reads and unpaired reads mapping to each locus. - 81 Date reçue / Received date 2024-08-20
  2. 4
    The method of any one of claims 1 to 3, wherein said duplex tags are not sequencing adaptors.
  3. 5
    The method of any one of claims 1 to 4, wherein the duplex tags are double-stranded tags that are Y-shaped with a hybridized portion at one end of the tag and a non-hybridized portion at the opposite end of the tag.
  4. 6
    The method of any one of claims 1 to 5, wherein the duplex tags contain molecular barcodes.
  5. 8
    The method of any one of claims 1 to 7, wherein reducing redundancy in said set of sequence reads comprises collapsing sequence reads produced from amplified products of an original polynucleotide molecule in said sample back to said original polynucleotide molecule.
  6. 11
    The method of any one of claims 8 to 10, further comprising determining a consensus sequence for said original polynucleotide molecule.
  7. 15
    The methods of any one of claims 12 to 14 wherein the sequence variant is a single nucleotide variant, an indel, a transversion, a translocation, an inversion, a deletion, a chromosomal structure alteration, a gene fusion, a chromosome fusion, a gene truncation, a gene amplification, a gene duplication, or a chromosomal lesion.
  8. 16
    A method, comprising:(a) providing a sample comprising a set of double-stranded polynucleotide molecules which are cell-free nucleic acid molecules, each double-stranded polynucleotide molecule including first and second complementary strands;(b) tagging said double-stranded polynucleotide molecules with a set of duplex tags, wherein each duplex tag differently tags said first and second complementary strands of a double-stranded polynucleotide molecule in said set;(c) sequencing at least some of said tagged strands to produce a set of sequence reads;-83Date reçue / Received date 2024-08-20 (d) reducing redundancy, tracking redundancy, or both in said set of sequence reads;(e) sorting sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to sequence reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in said set, and (ii) each unpaired read represents a first tagged strand having no second differently tag complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads;(f) determining quantitative measures of (i) said paired reads and (ii) said unpaired reads that map to each of one or more genetic loci;and (g) estimating with a programmed computer processor a quantitative measure of total double-stranded polynucleotide molecules in said set that map to each of said one or more genetic loci based on said quantitative measure of paired reads and unpaired reads mapping to each locus.
  9. 24
    A method, comprising:(a) from a sequencer, receiving into memory a set of sequence reads of polynucleotides tagged with duplex tags;(b) reducing redundancy, tracking redundancy, or both in said set of sequence reads;(c) sorting sequence reads in said set into paired reads and unpaired reads, wherein (i) each paired read corresponds to sequence reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule, and (ii) each unpaired read represents a first tagged strand having no second differently tag complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads;(d) determining quantitative measures of (i) said paired reads and (ii) said unpaired reads that map to each of one or more genetic loci;and -85Date reçue / Received date 2024-08-20 (e) estimating a quantitative measure of total double-stranded polynucleotide molecules that map to each of said one or more genetic loci based on said quantitative measure of paired reads and unpaired reads mapping to each locus. -86Date reçue / Received date 2024-08-20