EP3524694A1

Methods and systems for detecting genetic variants

Abstract

Disclosed herein in are methods and systems for determining genetic variants (e.g., copy number variation) in a polynucleotide sample. A method for determining copy number variations includes tagging double-stranded polynucleotides with duplex tags, sequencing polynucleotides from the sample and estimating total number of polynucleotides mapping to selected genetic loci. The estimate of total number of polynucleotides can involve estimating the number of double-stranded polynucleotides in the original sample for which no sequence reads are generated. This number can be generated using the number of polynucleotides for which reads for both complementary strands are detected and reads for which only one of the two complementary strands is detected.

EP3524694A1, drawing sheet 1
Sheet 1 of 13

Term

8.3 yearsto projected expiry

Projected expiry 24 December 2034, counted from filing; an application has no term until it is granted.

  1. Priority and filed
  2. Published
  3. Today
  4. Projected expiry

15 claims: 10 independent, 5 dependent

  1. 1
    A method for determining a quantitative measure indicative of a number of individual double-stranded deoxyribonucleic acid (DNA) molecules in a sample, comprising:(a) determining a quantitative measure of individual DNA molecules for which both strands are detected;(b) determining a quantitative measure of individual DNA molecules for which only one of the DNA strands is detected;(c) inferring from (a) and (b) above a quantitative measure of individual DNA molecules for which neither strand was detected;and (d) using (a)-(c) to determine the quantitative measure indicative of a number of individual double-stranded DNA molecules in the sample;wherein determining a quantitative measure of individual DNA molecules comprises: (i) tagging said DNA molecules with a set of duplex tags which differently tag complementary strands of a double-stranded DNA molecule in said sample to provide tagged strands;and (ii) sequencing at least some of said tagged strands to produce a set of sequence reads.
  2. 5
    The method of any one of claims 1 to 4, wherein the double-stranded DNA molecules comprise cfDNA.
  3. 6
    The method of any one of claims 1 to 5, wherein the duplex tags are double-stranded tags that are Y-shaped with a hybridized portion at one end of the tag and a non-hybridized portion is at the opposite end of the tag.
  4. 7
    The method of any one of claims 1 to 6, wherein the duplex tags contain molecular barcodes, optionally wherein the tagging occurs in a single reaction and results in greater than 50% of the DNA molecules being tagged at both ends.
  5. 9
    The method of any one of claims 7 to 8, wherein the method further comprises reducing or tracking redundancy in the sequence reads to determine consensus reads that are representative of single-strands of the original DNA, optionally wherein the method to reduce or track redundancy comprises comparing sequence reads having the same or similar molecular barcodes and the same or similar end of sequences.
  6. 10
    The method of any one of claims 7 to 9, wherein the method further comprises binning the sequence reads according to the molecular barcodes and sequence information, optionally wherein the binning is performed from at least one end of the original DNA to create bins of single stranded reads.
  7. 11
    The method of any one of claims 1 to 10, wherein the sample is derived from blood, plasma, serum, urine, saliva, mucosal excretions, sputum, stool, cerebral spinal fluid, skin, hair, sweat, and/or tears.
  8. 12
    The method of any one of claims 1 to 11, wherein the sample is derived from a subject suspected of having a disease, optionally wherein the disease is cancer.
  9. 13
    The method of any one of claims 1 to 12, wherein the method further comprises selectively enriching a subset of the tagged DNA, optionally wherein the selective enrichment is performed by hybridisation or amplification techniques.
  10. 15
    The method of any one of claims 1 to 14, wherein the method further comprises analyzing the nucleotide sequences with a programmed computer processor to identify one or more genetic alterations in the nucleotide sample of a subject, optionally wherein the one or more genetic alterations is selected from the list comprising base change(s), insertion(s), repeat(s), deletion(s), copy number variation(s), epigenetic modification(s), nucleosome binding site(s), copy number change(s) due to origin(s) of replication, and transversion(s).