Nova Patents
EP4585697A2

Methods for isolating cell-free dna

Abstract

Disclosed herein are methods for isolating DNA, such as cell-free DNA (cfDNA) or DNA from a tissue sample, e.g., in which the DNA is partitioned into hypermethylated and hypomethylated partitions. After differential tagging of the partitions, portions of the hypomethylated partition are pooled with the hypermethylated partition or pooled separately. Epigenetic and sequence-variable target regions are captured from the pool comprising DNA from the hypermethylated and hypomethylated partitions, and sequence-variable target regions are captured from the pool comprising DNA from the hypomethylated partition. This approach can reduce costs and/or bandwidth by limiting sequencing of epigenetic target regions from the hypomethylated partition, which may be less informative than other DNA.

EP4585697A2, drawing sheet 1
Sheet 1 of 9

Term

14.8 yearsto projected expiry

Projected expiry 29 July 2041, counted from filing; an application has no term until it is granted.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

16 claims: 13 independent, 3 dependent

  1. 1
    A method of isolating DNA from a sample, the method comprising:partitioning the DNA of the sample into a plurality of partitions, the plurality comprising at least a hypermethylated partition and a hypomethylated partition;preparing a first pool comprising at least a first portion of the DNA of the hypomethylated partition;preparing a second pool comprising at least a first portion of the DNA of the hypermethylated partition;capturing at least a first set of target regions from the first pool;and capturing at least a second set of target regions from the second pool, wherein the first set of target regions and the second set of target regions are not identical.
  2. 3
    A method of isolating DNA from a sample, the method comprising:partitioning the DNA of the sample into a plurality of partitions, the plurality comprising at least a hypermethylated partition and a hypomethylated partition;differentially tagging the DNA of the hypermethylated partition and the DNA of the hypomethylated partition;preparing a first pool comprising at least a first portion of the DNA of the hypomethylated partition;preparing a second pool comprising at least a first portion of the DNA of the hypermethylated partition;capturing at least a first set of target regions from the first pool, wherein the first set comprises sequence-variable target regions;and capturing a second plurality of sets of target regions from the second pool, wherein the second plurality comprises sequence-variable target regions and epigenetic target regions.
  3. 4
    The method of any one of claims 1-3, wherein:(i) capturing the first set of target regions from the first pool comprises contacting the DNA of the first pool with a first set of target-specific probes, optionally wherein the first set of target-specific probes comprises target-binding probes specific for sequence-variable target regions;(ii) capturing the second plurality of sets of target regions or second set of target regions from the second pool comprises contacting the DNA of the second pool with a second set of target-specific probes, optionally wherein the second set of target-specific probes comprises target-binding probes specific for sequence-variable target regions and/or target-binding probes specific for epigenetic target regions;(iii) the DNA comprises cell-free DNA (cfDNA);(iv) the first portion of the DNA of the hypomethylated partition comprises at least about 50% of the DNA of the hypomethylated partition;(v) the first portion of the DNA of the hypomethylated partition comprises about 50-95% of the DNA of the hypomethylated partition;and/or (vi) the first portion of the DNA of the hypomethylated partition comprises at least about 80% of the DNA of the hypomethylated partition.
  4. 5
    The method of any one of claims 1-4, wherein:(i) the second pool comprises a second portion of the DNA of the hypomethylated partition;(ii) the first portion of the DNA of the hypomethylated partition comprises a greater amount of DNA of the hypomethylated partition than the second portion of the DNA of the hypomethylated partition;(iii) the second portion of the DNA of the hypomethylated partition comprises less than or equal to about 50% of the DNA of the hypomethylated partition;and/or (iv) the second portion of the DNA of the hypomethylated partition comprises less than or equal to about 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% of the DNA of the hypomethylated partition.
  5. 6
    The method of any one of claims 1-4, wherein the first pool comprises substantially all of the DNA of the hypomethylated partition.
  6. 7
    The method of any one of claims 1-6, wherein:(i) the second portion comprises at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the DNA of the hypermethylated partition;(ii) the second pool comprises substantially all of the DNA of the hypermethylated partition;(iii) the plurality of partitions further comprises an intermediate partition, optionally wherein the second pool comprises at least a portion of the intermediate partition, such as wherein the second pool comprises substantially all of the intermediate partition;(iv) the second set of target regions or the second plurality of sets of target regions comprises a greater number of epigenetic target regions than the first set of target regions;(v) the first set of target regions comprises a greater amount of sequence-variable target regions than the second set of target regions or the second plurality of sets of target regions;(vi) the first set of target regions does not comprise epigenetic target regions;(vii) the epigenetic target region set comprises a hypermethylation variable target region set;and/or (viii) the epigenetic target region set comprise a fragmentation variable target region set or the first set of target regions comprises fragmentation-variable target regions, optionally wherein the fragmentation variable target region set comprises: (a) transcription start site regions;and/or (b) CTCF binding regions.
  7. 8
    The method of any one of claims 1-7, wherein:(I) the second plurality of sets of target regions comprises fragmentation-variable target regions;(II) at least one hypermethylation variable target region is captured from the second pool but not from the first pool, optionally wherein a plurality of hypermethylation variable target regions are captured from the second pool but not from the first pool;and/or (III) the method further comprises sequencing the first and second pluralities of sets of target regions, optionally wherein: (A) DNA molecules corresponding to the sequence-variable target region set are sequenced to a greater depth of sequencing than the cfDNA molecules corresponding to the epigenetic target region set;(B) the sequencing generates a plurality of sequence reads and the method further comprises mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads;(C) the sample is from a subject and the method further comprises determining the presence or absence of a cancer in the subject based at least in part on data generated by sequencing the first and second pluralities of sets of target regions;(D) the sample is from a subject and the method further comprises determining a likelihood that the subject has cancer based at least in part on data generated by sequencing the first and second pluralities of sets of target regions;(E) the method further comprises mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads and processing the mapped sequence reads corresponding to the sequence-variable target region set and to the epigenetic target region set to determine the likelihood that the subject has cancer;and/or (F) molecule counts are determined from the sequencing results of the hypermethylated and hypomethylated partitions, such as wherein a fraction of the DNA of the hypomethylated partition was included in the second pool, for example wherein molecule counts for epigenetic target regions in the hypomethylated partition are estimated by multiplication of observed molecule counts with a scaling factor, optionally wherein: (a) the scaling factor is the reciprocal of the fraction of the hypomethylated partition that was included in the second pool;(b) molecule counts for epigenetic target regions in the hypomethylated partition are estimated by using an anchor ratio determined based on control region frequencies;(c) molecule counts for epigenetic target regions in the hypomethylated partition are estimated by using an anchor ratio determined based on diversity levels;(d) molecule counts for epigenetic target regions in the hypomethylated partition are estimated by multiplication with a scaling factor determined (i) from a mean or median fold difference in frequency of epigenetic target regions in hypomethylated partition sequence data from samples in which an entire hypomethylated partition was sequenced versus samples in which only a portion of a hypomethylated partition was sequenced, or (ii) from a mean or median fold difference in frequency of epigenetic target regions in hypomethylated partition sequence data from a plurality of sets of sequence data from one or a plurality of samples, the sets of sequence data comprising sequence data in which a fraction of the hypomethylated partition was sequenced and sequence data in which the entire hypomethylated partition was sequenced;(e) molecule counts for epigenetic target regions in the hypomethylated partition are estimated by multiplication with a scaling factor determined using frequencies of epigenetic target regions for which probes were included in the capturing of sequence-variable target regions from the first pool;or (f) molecule counts for epigenetic target regions in the hypomethylated partition are estimated by using a relationship between reads and unique molecules to infer a molecule count that would have resulted from capturing epigenetic target regions from all of the hypomethylated partition.
  8. 9
    The method of any one of claims 1-8, wherein the test subject was previously diagnosed with a cancer and received one or more previous cancer treatments, optionally wherein the DNA is obtained at one or more preselected time points following the one or more previous cancer treatments.
  9. 12
    The method of any one of claims 9-11, further comprising determining a disease-free survival (DFS) period for the test subject based on the cancer recurrence score, optionally wherein the DFS period is 1 year, 2 years, 3, years, 4 years, 5 years, or 10 years.
  10. 13
    The method of any one of claims 10-12, wherein:(I) the set of sequence information comprises sequence-variable target region sequences, and determining the cancer recurrence score comprises determining at least a first subscore indicative of the amount of SNVs, insertions/deletions, CNVs and/or fusions present in sequence-variable target region sequences, optionally wherein a number of mutations in the sequence-variable target regions chosen from 1, 2, 3, 4, or 5 is sufficient for the first subscore to result in a cancer recurrence score classified as positive for cancer recurrence, optionally wherein the number of mutations is chosen from 1, 2, or 3;(II) the set of sequence information comprises epigenetic target region sequences, and determining the cancer recurrence score comprises determining a second subscore indicative of the amount of abnormal sequence reads in the epigenetic target region sequences, optionally wherein abnormal sequence reads comprise reads indicative of methylation of hypermethylation variable target sequences and/or reads indicative of abnormal fragmentation in fragmentation variable target regions, such as wherein a proportion of reads corresponding to the hypermethylation variable target region set and/or fragmentation variable target region set that indicate hypermethylation in the hypermethylation variable target region set and/or abnormal fragmentation in the fragmentation variable target region set greater than or equal to a value in the range of 0.001%-10% is sufficient for the second subscore to be classified as positive for cancer recurrence, for example wherein the range is: (a) 0.001%-1% or 0.005%-1%;(b) 0.01%-5% or 0.01%-2%;or (c) 0.01%-1%;(III) the method further comprises determining a fraction of tumor DNA from the fraction of reads in the plurality of sequence reads that indicate one or more features indicative of origination from a tumor cell, optionally wherein: (a) the one or more features indicative of origination from a tumor cell comprise one or more of alterations in a sequence-variable target region, hypermethylation of a hypermethylation variable target region, and abnormal fragmentation of a fragmentation variable target region;(b) the method further comprises determining a cancer recurrence score based at least in part on the fraction of tumor DNA, wherein a fraction of tumor DNA greater than or equal to a predetermined value in the range of 10 -11 to 1 or 10 -10 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence, such as wherein: (i) a fraction of tumor DNA greater than or equal to a predetermined value in the range of 10 -10 to 10 -9 , 10 -9 to 10 -8 , 10 -8 to 10 -7 , 10 -7 to 10 -6 , 10 -6 to 10 -5 , 10 -5 to 10 -4 , 10 -4 to 10 -3 , 10 -3 to 10 -2 , or 10 -2 to 10 -1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence;and/or (ii) the predetermined value is in the range of 10 -8 to 10 -6 or is 10 -7 ;and/or (c) the fraction of tumor DNA is determined as greater than or equal to the predetermined value if the cumulative probability that the fraction of tumor DNA is greater than or equal to the predetermined value is at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995, or 0.999, optionally wherein the cumulative probability is: (i) at least 0.95;or (ii) in the range of 0.98-0.995 or is 0.99;and/or (IV) the set of sequence information comprises sequence-variable target region sequences and epigenetic target region sequences, and determining the cancer recurrence score comprises determining a first subscore indicative of the amount of SNVs, insertions/deletions, CNVs and/or fusions present in sequence-variable target region sequences and a second subscore indicative of the amount of abnormal sequence reads in epigenetic target region sequences, and combining the first and second subscores to provide the cancer recurrence score, optionally wherein combining the first and second subscores comprises applying a threshold to each subscore independently (e.g., greater than a predetermined number of mutations (e.g., > 1) in sequence-variable target regions, and greater than a predetermined fraction of abnormal (e.g., tumor) reads in epigenetic target regions), or training a machine learning classifier to determine status based on a plurality of positive and negative training samples, such as wherein a value for the combined score in the range of -4 to 2 or -3 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence.
  11. 14
    The method of any one of claims 9-13, wherein:(i) the one or more preselected timepoints is selected from the following group consisting of 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, 1.5 years, 2 year, 3 years, 4 and 5 years after administration of the one or more previous cancer treatments;(ii) the cancer is colorectal cancer;(iii) the one or more previous cancer treatments comprise surgery;(iv) the one or more previous cancer treatments comprise administration of a therapeutic composition;and/or (v) the one or more previous cancer treatments comprise chemotherapy.
  12. 15
    The method of any one of claims 1-14, wherein:(i) DNA molecules corresponding to the sequence-variable target region set are captured from the second pool with a greater capture yield than DNA molecules corresponding to the epigenetic target region set. optionally wherein the captured DNA molecules of the sequence-variable target region set are: (a) sequenced to at least a 2-fold greater depth of sequencing than the captured DNA molecules of the epigenetic target region set;(b) sequenced to at least a 3-fold greater depth of sequencing than the captured DNA molecules of the epigenetic target region set;(c) sequenced to a 4-10-fold greater depth of sequencing than the captured DNA molecules of the epigenetic target region set;or (d) sequenced to a 4-100-fold greater depth of sequencing than the captured DNA molecules of the epigenetic target region set;(ii) the sequence-variable target regions are sequenced to at least 1000X coverage, optionally wherein the sequence-variable target regions are sequenced to a greater amount of coverage than the epigenetic target regions;(iii) the sequence-variable target regions are sequenced to an amount of coverage in the range of 1000X-20,000X, optionally wherein the sequence-variable target regions are sequenced to a greater amount of coverage than the epigenetic target regions;(iv) the epigenetic target regions are sequenced to at least 1000X coverage, optionally wherein the sequence-variable target regions are sequenced to a greater amount of coverage than the epigenetic target regions;(v) the epigenetic target regions are sequenced to an amount of coverage in the range of 1000X-10,000X, optionally wherein the sequence-variable target regions are sequenced to a greater amount of coverage than the epigenetic target regions;and/or (vi) the first set of target regions are pooled with the second set of target regions or the second plurality of sets of target regions before sequencing;optionally wherein the captured DNA molecules of the sequence-variable target region set and the captured DNA molecules of the epigenetic target region set are sequenced in the same sequencing cell.
  13. 16
    The method of any one of claims 1-15, wherein:(i) the DNA is amplified before capture, optionally wherein the method further comprises ligating barcode-containing adapters to the DNA when or before the DNA is amplified;(ii) capturing the second set of target regions of DNA or second plurality of sets of target regions of DNA comprises contacting the DNA with target-binding probes specific for a sequence-variable target region set and target-binding probes specific for an epigenetic target region set, optionally wherein target-binding probes specific for the sequence-variable target region set are: (a) present in a higher concentration than the target-binding probes specific for the epigenetic target region set;(b) present in at least a 2-fold higher concentration than the target-binding probes specific for the epigenetic target region set;or (c) present in at least a 4-fold or 5-fold higher concentration than the target-binding probes specific for the epigenetic target region set;optionally wherein target-binding probes specific for the sequence-variable target region set have a higher target binding affinity than the target-binding probes specific for the epigenetic target region set;(iii) the epigenetic target region set has a footprint which is at least 2-fold greater than the size of the sequence-variable target region set, optionally wherein the footprint of the epigenetic target region set is at least 10-fold greater than the size of the sequence-variable target region set;(iv) the sequence-variable target region set has a footprint of at least 25 kB or 50 kB;(v) the DNA obtained from the test subject is partitioned into at least 2 fractions on the basis of methylation level, and the subsequent steps of the method are performed on each fraction, optionally wherein: (a) the partitioning step comprises contacting the collected DNA with a methyl binding reagent immobilized on a solid support, optionally wherein the methyl binding reagent comprises a methyl binding domain or methyl binding protein;and/or (b) the at least 2 fractions comprise a hypermethylated fraction and a hypomethylated fraction, and the method further comprises differentially tagging the hypermethylated fraction and the hypomethylated fraction or separately sequencing the hypermethylated fraction and the hypomethylated fraction, such as wherein the hypermethylated fraction and the hypomethylated fraction are differentially tagged and the method further comprises pooling the differentially tagged hypermethylated and hypomethylated fractions before a sequencing step;(vi) the method further comprises determining whether DNA molecules corresponding to the sequence-variable target region set comprise cancer-associated mutations;(vii) the method further comprises determining whether DNA molecules corresponding to the epigenetic target region set comprise or indicate cancer-associated epigenetic modifications or copy number variations (e.g., focal amplifications), optionally wherein the method comprises determining whether DNA molecules corresponding to the epigenetic target region set comprise or indicate cancer-associated epigenetic modifications and copy number variations (e.g., focal amplifications), optionally wherein the cancer-associated epigenetic modifications comprise: (a) hypermethylation in one or more hypermethylation variable target regions;(b) one or more perturbations of CTCF binding;and/or (c) one or more perturbations of transcription start sites;and/or (viii) the captured sets of DNA molecules are sequenced using high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore-based sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), massively-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Sanger sequencing, Maxam-Gilbert sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or a Nanopore platform.